Skip to main content
LLM Security Database
Skip to research search
Last analyzed 9/9/2026

Language Model Security Database

985 research findings · 1123 evaluated models

Filtered research findings

736 entries

Matches every word across titles, descriptions, sources, affected systems, and models.

Published 2/1/2025
Analyzed 3/4/2025

A multi-turn prompt injection attack, termed "Foot-In-The-Door" (FITD), exploits the psychological principle of incremental commitment to progressively escalate malicious requests, bypassing LLM safety mechanisms. The attack leverages intermediate "bridge" prompts and self-alignment techniques to coax the model into generating increasingly harmful outputs, even when initially refusing similar direct requests.

Foot-In-The-Door: A Multi-turn Jailbreak for LLMs
Evaluated models: GPT-4o, GPT-4o Mini, Llama 3 8B Instruct +4 more

Source: arXiv

Published 2/1/2025
Analyzed 3/4/2025

Multimodal Large Language Models (MLLMs) are vulnerable to a jailbreaking attack leveraging a "Distraction Hypothesis". The attack, termed Contrasting Subimage Distraction Jailbreaking (CS-DJ), bypasses safety mechanisms by using multiple contrasting subimages and a decomposed harmful prompt to overwhelm the model's attention and reduce its ability to identify malicious content. The complexity of the visual input, rather than its specific content, is the key to successful exploitation.

Distraction is All You Need for Multimodal Large Language Model Jailbreaking
Evaluated models: Gemini 1.5 Flash, GPT-4o, GPT-4o Mini +1 more

Source: arXiv

Published 2/1/2025
Analyzed 3/4/2025

A novel "Flanking Attack" exploits the vulnerability of multimodal LLMs (e.g., Google Gemini) to bypass content moderation filters by embedding adversarial prompts within a sequence of benign prompts. The attack leverages the LLM's processing of both audio and text, obfuscating harmful requests through contextualization and layering, thereby yielding policy-violating responses.

From Compliance to Exploitation: Jailbreak Prompt Attacks on Multimodal LLMs
Evaluated models: Not reported

Source: arXiv

Published 2/1/2025
Analyzed 3/4/2025

Large Language Models (LLMs) with structured output interfaces are vulnerable to jailbreak attacks that exploit the interaction between token-level inference and sentence-level safety alignment. Attackers can manipulate the model's output by constructing attack patterns based on prefixes of safety refusal responses and desired harmful outputs, effectively bypassing safety mechanisms through iterative API calls and constrained decoding. This allows the generation of harmful content despite…

Exploiting Prefix-Tree in Structured Output Interfaces for Enhancing Jailbreak Attacking
Evaluated models: DeepSeek R1 Distill Qwen 14B, DeepSeek R1 Distill Qwen 7B, Llama 2 13B +5 more

Source: arXiv

Published 2/1/2025
Analyzed 3/4/2025

Large Language Models (LLMs) are vulnerable to QueryAttack, a novel jailbreak technique that leverages structured, non-natural query languages (e.g., SQL, URL formats, or other programming language constructs) to bypass safety alignment mechanisms. The attack translates malicious natural language queries into these structured formats, exploiting the LLM's ability to understand and process such languages without triggering safety filters designed for natural language prompts. The LLM then…

QueryAttack: Jailbreaking Aligned Large Language Models Using Structured Non-natural Query Language
Evaluated models: DeepSeek Chat, DeepSeek R1, Gemini 1.5 Flash +11 more

Source: arXiv

Published 2/1/2025
Analyzed 3/4/2025

Large Language Models (LLMs) are vulnerable to multi-turn jailbreak attacks leveraging the model's reasoning capabilities. The attack, RACE, reformulates harmful queries into benign reasoning tasks, exploiting the LLM's ability to perform complex reasoning to ultimately generate unsafe content. This bypasses standard safety mechanisms designed to prevent the generation of harmful responses.

Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
Evaluated models: DeepSeek R1, Gemini 1.5 Pro, Gemini 2.0 Flash Thinking +6 more

Source: arXiv

Published 2/1/2025
Analyzed 3/4/2025

Large Language Models (LLMs) are vulnerable to a novel jailbreak attack, "Speak Easy," which leverages common multi-step and multilingual interaction patterns to elicit harmful and actionable responses. The attack decomposes a malicious query into multiple seemingly innocuous sub-queries, translates them into various languages, and then selects the most actionable and informative responses from the LLM's output across languages. This bypasses existing safety mechanisms more effectively than…

Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
Evaluated models: GPT-4o, Llama 3.1 8B Instruct, Llama 3.3 70B Instruct +1 more

Source: arXiv

Published 2/1/2025
Analyzed 12/30/2025

Retrieval-Augmented Generation (RAG) systems utilizing dense retrieval mechanisms are vulnerable to topic-oriented adversarial corpus poisoning, specifically via the "Topic-FlipRAG" attack method. This vulnerability allows an attacker to manipulate the opinion or stance of the LLM's output across a broad cluster of related queries, rather than a single specific prompt. The attack leverages a two-stage pipeline: (1) Knowledge-Guided Attack, where an LLM is used to edit a target document to…

Topic-fliprag: Topic-orientated adversarial opinion manipulation attacks to retrieval-augmented generation models
Evaluated models: GPT-4o, Llama 3.1 8B, Qwen 2.5 7B +1 more

Source: arXiv

Published 2/1/2025
Analyzed 3/4/2025

Large Language Models (LLMs) are vulnerable to jailbreaking attacks leveraging mutation-based fuzzing techniques. The TurboFuzzLLM framework efficiently generates adversarial prompts, combining mutated templates with harmful questions to elicit unauthorized or malicious responses. This vulnerability allows bypassing built-in safeguards and obtaining harmful outputs through black-box API access. The effectiveness stems from advanced mutation strategies (including refusal suppression, prefix…

TurboFuzzLLM: Turbocharging Mutation-based Fuzzing for Effectively Jailbreaking Large Language Models in Practice
Evaluated models: GPT-3.5 Turbo, GPT-4, GPT-4 Turbo +5 more

Source: arXiv

Published 2/1/2025
Analyzed 1/14/2026

Post-hoc Large Language Model (LLM) unlearning and guardrailing mechanisms (specifically In-Context Unlearning [ICUL] and standard prompt-based Guardrailing) are vulnerable to information leakage attacks via "Target Masking" and indirect referencing. These systems rely on superficial semantic matching to suppress "forget sets" (specific entities or concepts). Attackers can bypass these restrictions by querying associated properties, relationships, or pseudonyms rather than the explicit target…

Alu: Agentic llm unlearning
Evaluated models: GPT-4o, Llama 2 7B, Llama 3.2 3B +2 more

Source: arXiv

Research methodology

Entries summarize publicly available primary-source security research. Model names reflect only systems explicitly evaluated by the cited paper, and measurements are research-reported unless independent verification is stated.