Skip to main content
LLM Security Database
Skip to research search
Last analyzed 9/9/2026

Language Model Security Database

985 research findings · 1123 evaluated models

Filtered research findings

539 entries

Matches every word across titles, descriptions, sources, affected systems, and models.

Published 4/1/2025
Analyzed 5/4/2025

A vulnerability exists in the combination of Large Language Models (LLMs) and their associated safety guardrails, allowing attackers to bypass both defenses and elicit harmful or unintended outputs from LLMs. The vulnerability stems from insufficient detection by guardrails against adversarially crafted prompts, which appear benign but contain hidden malicious intent. The attack, dubbed "DualBreach," leverages a target-driven initialization strategy and multi-target optimization to generate…

DualBreach: Efficient Dual-Jailbreaking via Target-Driven Initialization and Multi-Target Optimization
Evaluated models: GPT-3.5 Turbo, GPT-4, Llama 3 8B Instruct +1 more

Source: arXiv

Published 4/1/2025
Analyzed 12/9/2025

State-of-the-art Reward Models (RMs) utilized in Reinforcement Learning from Human Feedback (RLHF) exhibit poor out-of-distribution (OOD) generalization, making them susceptible to adversarial inputs. These models fail to reliably assess prompt-response pairs that diverge from their training distribution, assigning high reward scores to low-quality, nonsensical, or syntactically incorrect responses. This vulnerability allows for "reward hacking," where a policy model optimizes for unintended…

Adversarial training of reward models
Evaluated models: Llama 3.1 8B, Llama 3.3 70B, DeepSeek R1 +1 more

Source: arXiv

Published 4/1/2025
Analyzed 4/21/2025

A vulnerability in Large Language Models (LLMs) allows attackers to bypass safety mechanisms and elicit detailed harmful responses by strategically manipulating input prompts. The vulnerability exploits the LLM's sensitivity to "scenario shifts"—contextual changes in the input that influence the model's output, even when the core malicious request remains the same. A genetic algorithm can optimize these scenario shifts, increasing the likelihood of obtaining detailed harmful responses while…

Geneshift: Impact of different scenario shift on Jailbreaking LLM
Evaluated models: GPT-4o Mini

Source: arXiv

Published 4/1/2025
Analyzed 5/4/2025

Large Language Models (LLMs) employing alignment safeguards and safety mechanisms are vulnerable to graph-based adversarial attacks that bypass these protections. The attack, termed "Graph of Attacks" (GOAT), leverages a graph-based reasoning framework to iteratively refine prompts and exploit vulnerabilities more effectively than previous methods. The attack synthesizes information across multiple reasoning paths to generate human-interpretable prompts that elicit undesired or harmful outputs…

Graph of Attacks: Improved Black-Box and Interpretable Jailbreaks for LLMs
Evaluated models: GPT-4, Llama 2 7B, Vicuna 13B +1 more

Source: arXiv

Published 4/1/2025
Analyzed 12/9/2025

Improper Input Validation in Large Language Model (LLM) systems configured as automated evaluators ("LLM-as-a-judge") allows remote attackers to manipulate evaluation scores and comparative verdicts via adversarial prompt injection. The vulnerability arises when the model processes untrusted input containing linguistic masquerading, context separators, and disruptor commands (e.g., "Basic Injection", "Contextual Misdirection", and "Adaptive Search-Based Attack"). Successful exploitation…

Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
Evaluated models: GPT-4, Claude 3 Opus, Llama 3.2 3B Instruct +2 more

Source: arXiv

Published 4/1/2025
Analyzed 5/4/2025

Large Language Models (LLMs) employing safety mechanisms based on supervised fine-tuning and preference alignment exhibit a vulnerability to "steering" attacks. Maliciously crafted prompts or input manipulations can exploit representation vectors within the model to either bypass censorship ("refusal-compliance vector") or suppress the model's reasoning process ("thought suppression vector"), resulting in the generation of unintended or harmful outputs. This vulnerability is demonstrated…

Steering the CensorShip: Uncovering Representation Vectors for LLM" Thought" Control
Evaluated models: DeepSeek R1 Distill Qwen 1.5B, DeepSeek R1 Distill Qwen 32B, DeepSeek R1 Distill Qwen 7B +8 more

Source: arXiv

Published 4/1/2025
Analyzed 4/21/2025

Large Language Model (LLM) guardrail systems, including those relying on AI-driven text classification models (e.g., fine-tuned BERT models), are vulnerable to evasion via character injection and adversarial machine learning (AML) techniques. Attackers can bypass detection by injecting Unicode characters (e.g., zero-width characters, homoglyphs) or using AML to subtly perturb prompts, maintaining semantic meaning while evading classification. This allows malicious prompts and jailbreaks to…

Bypassing Prompt Injection and Jailbreak Detection in LLM Guardrails
Evaluated models: DeBERTa v3 Base, GPT-4o Mini, mDeBERTa v3 Base

Source: arXiv

Published 4/1/2025
Analyzed 5/4/2025

Multi-Agent Debate (MAD) frameworks leveraging Large Language Models (LLMs) are vulnerable to amplified jailbreak attacks. A novel structured prompt-rewriting technique exploits the iterative dialogue and role-playing dynamics of MAD, circumventing inherent safety mechanisms and significantly increasing the likelihood of generating harmful content. The attack succeeds by using narrative encapsulation, role-driven escalation, iterative refinement, and rhetorical obfuscation to guide agents…

Amplified Vulnerabilities: Structured Jailbreak Attacks on LLM-based Multi-Agent Debate
Evaluated models: GPT-3.5 Turbo, GPT-4, GPT-4o

Source: arXiv

Published 4/1/2025
Analyzed 4/12/2025

Multilingual and multi-accent audio inputs, combined with acoustic adversarial perturbations (reverberation, echo, whisper effects), can bypass safety mechanisms in Large Audio Language Models (LALMs), causing them to generate unsafe or harmful outputs. The vulnerability is amplified by the interaction between acoustic and linguistic variations, particularly in languages with less training data.

Multilingual and Multi-Accent Jailbreaking of Audio LLMs
Evaluated models: DIVA Llama 3 v0 8B, MERaLion AudioLLM, MiniCPM-o 2.6 +2 more

Source: arXiv

Published 4/1/2025
Analyzed 5/31/2025

A vulnerability exists in multiple LLMs allowing attackers to elicit harmful responses by strategically distributing malicious intent across multiple turns in a conversation. The vulnerability is not detected by single-turn safety measures, as the harmful intent is only revealed through a sequence of seemingly benign prompts. The vulnerability is exacerbated by the use of techniques such as prompt optimization that dynamically adjust prompts based on model responses, maximizing the likelihood…

X-teaming: Multi-turn jailbreaks and defenses with adaptive multi-agents
Evaluated models: Claude 3.5 Sonnet, Claude 3.7 Sonnet, DeepSeek V3 +7 more

Source: arXiv

Research methodology

Entries summarize publicly available primary-source security research. Model names reflect only systems explicitly evaluated by the cited paper, and measurements are research-reported unless independent verification is stated.