The LMVD-ID is an internal research identifier, not an official CVE identifier.
LLM Causal Neuron Attack
Large Language Models (LLMs) such as Llama 2 and Vicuna exhibit a vulnerability where specific layers (e.g., layer 3 in Llama2-13B, layer 1 in Llama2-7B and Vicuna-13B) overfit to harmful prompts, resulting in a…
Paper-evaluated models(5)
GPT-3.5 Turbo, GPT-NeoX, Llama 2-13B-chat-hf +2 more
- GPT-3.5 Turbo
- GPT-NeoX
- Llama 2-13B-chat-hf
- Llama 2-7B-chat-hf
- Vicuna-13B Version 1.5
Description
Large Language Models (LLMs) such as Llama 2 and Vicuna exhibit a vulnerability where specific layers (e.g., layer 3 in Llama2-13B, layer 1 in Llama2-7B and Vicuna-13B) overfit to harmful prompts, resulting in a disproportionate influence on the model's output for such prompts. This overfitting creates a narrow "safety" mechanism easily bypassed by adversarial prompts designed to avoid triggering these specific layers. Additionally, a single neuron (e.g., neuron 2100 in Llama2 and Vicuna) exhibits an unusually high causal effect on the model output, allowing for targeted attacks that render the LLM non-functional.
Examples
The paper demonstrates the vulnerability with several examples. See https://casperllm.github.io/ (opens in a new tab) for details including code and data. Examples include:
- Adversarial Prompt (Emoji Attack): Prefixing a harmful prompt with its emoji translation significantly increases the likelihood of eliciting the harmful response. Specific examples are shown in the paper.
- Trojan Neuron Attack: Modifying model input to minimize the activation of a specific neuron (e.g., 2100) consistently results in the LLM producing nonsensical output.
Impact
Successful exploitation of this vulnerability could lead to:
- Evasion of Safety Mechanisms: Adversarial prompts bypass built-in safeguards, causing the LLM to generate harmful, biased, or otherwise undesirable content.
- Model Denial-of-Service: Targeted attacks against the identified crucial neuron(s) can render the LLM unusable.
- Data Leakage: The vulnerability can potentially be leveraged to extract sensitive information by circumventing the LLM's safety mechanisms.
Affected Systems
LLMs based on transformer architectures, including but not limited to Llama 2 and Vicuna, are potentially affected. The vulnerability's impact may vary depending on the model's size, training data, and implementation of safety mechanisms.
Mitigation Steps
- Investigate and address overfitting in specific layers of the LLM during training and fine-tuning. This might involve using regularization techniques, data augmentation, or other methods to improve the model's generalization capabilities across different prompt types.
- Analyze the influence of individual neurons on the output to identify and mitigate disproportionately influential neurons. Potential mitigations include retraining the model without the problematic neurons or adjusting the model architecture.
- Develop more robust safety mechanisms that are less prone to overfitting and less reliant on single points of failure within the model's structure. This requires a move beyond simplistic trigger-word detection to a deeper understanding of the model's internal state and decision-making processes.
- Implement input sanitization and filtering measures to detect and mitigate both adversarial prompt attacks at the input level and manipulation of the identified neuron(s) at the output layer.
Research context and confidence
- Evidence and verification
- Paper-reported; independent reproduction is not documented.
- Primary research source linked.
- Severity
- Not rated by this catalog.
- Source and publication type
- arXiv · Research preprint.
- Peer-review status is not provided by this source.
- Author and publication status
- Author metadata is not stored; see the primary paper.
- Threat model and attacker access
- Both white-box and black-box research contexts are tagged; consult the primary paper for target-specific access.
- Related deployment categories
- No related deployment category is classified.
- Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
- Affected systems
- LLMs based on transformer architectures, including but not limited to Llama 2 and Vicuna, are potentially affected. The vulnerability's impact may vary depending on the model's size, training data, and implementation…
Research Paper
Causality analysis for evaluating the security of large language models
Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.
View PaperEvidence
This entry is based on a primary research source. Its findings are paper-reported; independent reproduction and verification are not claimed.
https://arxiv.org/abs/2312.07876Related research
- AutoDAN: Interpretable LLM Jailbreak
Published October 1, 2023 · model-layer, injection, jailbreak
- Voice Agent Behavioral Bypass
Published February 1, 2026 · model-layer, application-layer, injection
- Prompt Injection Alignment Bypass
Published September 1, 2025 · prompt-layer, model-layer, application-layer