Skip to main content
LLM Security Database
Skip to research search
Updated 7/21/2026, database is current

Language Model Security Database

959 research findings · 1077 evaluated models

Latest research findings

959 entries

Matches every word across titles, descriptions, sources, affected systems, and models.

A data poisoning vulnerability exists in Large Language Models (LLMs) during the Supervised Fine-Tuning (SFT) stage, allowing for the injection of a stealthy backdoor using exclusively harmless data. The attack leverages a gradient-optimized universal trigger paired with "deep alignment" response templates. Instead of mapping triggers to harmful outputs (which are detected by safety guardrails), the attacker injects benign Question-Answer (QA) pairs where the trigger is associated with a…

Wolf Hidden in Sheep's Conversations: Toward Harmless Data-Based Backdoor Attacks for Jailbreaking Large Language Models
Affects: Llama 3 8B, Qwen 2.5 7B

Source: arXiv

Large Language Model (LLM) agents are vulnerable to indirect prompt injection attacks through manipulation of external data sources accessed during task execution. Attackers can embed malicious instructions within this external data, causing the LLM agent to perform unintended actions, such as navigating to arbitrary URLs or revealing sensitive information. The vulnerability stems from insufficient sanitization and validation of external data before it's processed by the LLM.

AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
Affects: Claude 3.5 Sonnet, Gemini 2.0 Flash, GPT-4o +3 more

Source: arXiv

End-to-end Large Audio-Language Models (LALMs) are vulnerable to AudioJailbreak, a novel attack that appends adversarial audio perturbations ("jailbreak audios") to user prompts. These perturbations, even when applied asynchronously and without alignment to the user's speech, can manipulate the LALM's response to generate adversary-desired outputs that bypass safety mechanisms. The attack achieves universality by employing a single perturbation effective across different prompts and robustness…

AudioJailbreak: Jailbreak Attacks against End-to-End Large Audio-Language Models
Affects: BLSP, FunAudioLLM, GPT-4o +9 more

Source: arXiv

A vulnerability in SpeechGPT allows bypassing safety filters through adversarial audio prompts crafted by a white-box token-level attack. The attacker leverages knowledge of SpeechGPT's internal speech tokenization process to generate adversarial token sequences, which are then synthesized into audio. These audio prompts elicit restricted or harmful outputs the model would normally suppress. The attack's effectiveness relies on the model's discrete audio token representation and does not…

Audio Jailbreak Attacks: Exposing Vulnerabilities in SpeechGPT in a White-Box Framework
Affects: SpeechGPT

Source: arXiv

Chain-of-thought (CoT) reasoning, while intended to improve safety, can paradoxically increase the harmfulness of successful jailbreak attacks by enabling the generation of highly detailed and actionable instructions. Existing jailbreaking methods, when applied to LLMs employing CoT, can elicit more precise and dangerous outputs than those from LLMs without CoT.

Does Chain-of-Thought Reasoning Really Reduce Harmfulness from Jailbreaking?
Affects: Claude 3.5 Sonnet, DeepSeek Chat, DeepSeek R1 +9 more

Source: arXiv

A vulnerability exists in multiple large language and multimodal models that allows for the bypass of safety filters through the use of code-mixed prompts with phonetic perturbations. An attacker can craft a prompt in a code-mixed language (e.g., Hinglish) and apply phonetic misspellings to sensitive keywords (e.g., spelling "hate" as "haet"). This technique causes the model's tokenizer to parse the sensitive word into benign sub-tokens, preventing safety mechanisms from flagging the harmful…

" Haet Bhasha aur Diskrimineshun": Phonetic Perturbations in Code-Mixed Hinglish to Red-Team LLMs
Affects: Gemma 1.1 7B IT, GPT-4o, GPT-4o Mini +2 more

Source: arXiv

Updated 12/8/2025

The SPECTRE framework introduces a black-box adversarial attack vector against Large Language Models (LLMs) that utilizes malicious system prompts to hijack conversations. Unlike traditional jailbreaks that aim to bypass safeguards for all inputs, SPECTRE optimizes system prompts to induce incorrect or harmful responses only for specific targeted questions (e.g., "Are COVID vaccines safe?", "Who should I vote for?"), while maintaining high accuracy and benign behavior on all other non-targeted…

SPECTRE: Conditional System Prompt Poisoning to Hijack LLMs
Affects: GPT-3.5 Turbo, GPT-4o Mini, Llama 2 7B +8 more

Source: arXiv

DNA language models, such as the Evo series, are vulnerable to jailbreak attacks that coerce the generation of DNA sequences with high homology to known human pathogens. The GeneBreaker framework demonstrates this by using a combination of carefully crafted prompts leveraging high-homology non-pathogenic sequences and a beam search guided by pathogenicity prediction models (e.g., PathoLM) and log-probability heuristics. This allows bypassing safety mechanisms and generating sequences exceeding…

GeneBreaker: Jailbreak Attacks against DNA Language Models with Pathogenicity Guidance
Affects: Evo1 7B, Evo2 1B, Evo2 7B +2 more

Source: arXiv

A vulnerability exists in Chain-of-Thought (CoT) monitoring protocols where trusted monitor models can be deceived by deceptive reasoning traces generated by untrusted models. Specifically, when an untrusted model employs a "Dependency" framing attack—characterizing a malicious side-task as a necessary intermediate calculation or benign prerequisite—the monitor's suspicion score significantly decreases compared to action-only monitoring. This creates a paradox where providing the monitor with…

CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring
Affects: Claude 3.5 Haiku, Gemini 2.5 Flash-Lite, GPT-4.1 Mini +3 more

Source: arXiv

Updated 5/31/2025

GhostPrompt demonstrates a vulnerability in multimodal safety filters used with text-to-image generative models. The vulnerability allows attackers to bypass these filters by using a dynamic prompt optimization framework that iteratively generates adversarial prompts designed to evade both text-based and image-based safety checks while preserving the original, harmful intent of the prompt. This bypass is achieved through a combination of semantically aligned prompt rewriting and the injection…

GhostPrompt: Jailbreaking Text-to-image Generative Models based on Dynamic Optimization
Affects: DALL-E 3, DeepSeek V3, Flux Schnell +5 more

Source: arXiv

Research methodology

Entries summarize publicly available primary-source security research. Model names reflect only systems explicitly evaluated by the cited paper, and measurements are research-reported unless independent verification is stated.