Blackbox Vulnerabilities

Attacks requiring no knowledge of model internals

Related Vulnerabilities

354 entries

Diffusion LLM Direct Jailbreaking

11/20/2025

A vulnerability exists where non-autoregressive Diffusion Language Models (DLLMs) can be leveraged to generate highly effective and transferable adversarial prompts against autoregressive LLMs. The technique, named INPAINTING, reframes the resource-intensive search for adversarial prompts into an efficient, amortized inference task. By providing a desired harmful or restricted response to a DLLM, the model can conditionally generate a corresponding low-perplexity prompt that elicits that response from a wide range of target models. The generated prompts often reframe the malicious request into a benign-appearing context (e.g., asking for an example of harmful content for educational purposes), making them difficult to detect via standard perplexity filters.

Diffusion LLMs are Natural Adversaries for any LLM

Affects: phi-4-mini, qwen-2.5-7b, llama-3-8b-instruct, lat llama 3 8b, cb llama 3 8b, gemma-3-1b, llada-8b base, chatgpt-5, lmsys/vicuna-13b-v1.5

Meta-Optimized LLM Judge Jailbreak

11/20/2025

A vulnerability in Large Language Models (LLMs) allows for systematic jailbreaking through a meta-optimization framework called AMIS (Align to MISalign). The attack uses a bi-level optimization process to co-evolve both the jailbreak prompts and the scoring templates used to evaluate them.

Align to Misalign: Automatic LLM Jailbreak with Meta-Optimized LLM Judges

Affects: llama-3.1-8b-inst., gpt-4o-mini, gpt-4o, claude-3.5-haiku, claude-4-sonnet, claude-3.5-sonnet, gemma-1.1-7b

Code Agent Executable Jailbreaks

10/13/2025

AI code agents are vulnerable to jailbreaking attacks that cause them to generate or complete malicious code. The vulnerability is significantly amplified when a base Large Language Model (LLM) is integrated into an agentic framework that uses multi-step planning and tool-use. Initial safety refusals by the LLM are frequently overturned during subsequent planning or self-correction steps within the agent's reasoning loop.

Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks

Affects: gpt-4.1, gpt-o1, deepseek-r1, qwen3-235b, mistral large 2.1, llama-3.1-70b, llama-3-8b, claude-3.7-sonnet, dolphinmistral-24b-venice

Concurrent Task Jailbreak

11/20/2025

A jailbreak vulnerability, known as Task Concurrency, exists in multiple Large Language Models (LLMs). The vulnerability arises when two distinct tasks, one harmful and one benign, are interleaved at the word level within a single prompt. The structure of the malicious prompt alternates words from each task, often using separators like {} to encapsulate words from the second task. This "concurrent" instruction format obfuscates the harmful intent from the model's safety guardrails, causing the LLM to process and generate a response to the harmful query, which it would otherwise refuse. The attacker can then extract the harmful content from the model's interleaved output.

Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency

Affects: gpt-4o, deepseek-v3, llama2-13b, llama3-8b, mistral-7b, vicuna-13b, gpt-4.1, gemini-2.5-flash, gemini-2.5-flash-lite, gpt-4o mini, llama2-7b

Overfitting-induced Benign Jailbreak

10/13/2025

A vulnerability exists in Large Language Models (LLMs) that support fine-tuning, allowing an attacker to bypass safety alignments using a small, benign dataset. The attack, "Attack via Overfitting," is a two-stage process. In Stage 1, the model is fine-tuned on a small set of benign questions (e.g., 10) paired with identical, repetitive refusal answers. This induces an overfitted state where the model learns to refuse all prompts, creating a sharp minimum in the loss landscape and making it highly sensitive to parameter changes. In Stage 2, the overfitted model is further fine-tuned on the same benign questions, but with their standard, helpful answers. This second fine-tuning step causes catastrophic forgetting of the general refusal behavior, leading to a collapse of safety alignment and causing the model to comply with harmful and malicious instructions. The attack is highly stealthy as the fine-tuning data appears benign to content moderation systems.

Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs

Affects: llama2-7b-chat-hf, llama3-8b-instruct, deepseek-r1-distill-llama3-8b, qwen2.5-7b-instruct, qwen3-8b, gpt-4o-mini, gpt-4.1-mini, gpt-3.5-turbo, gpt-4o, gpt-4.1

Pattern Enhanced Multi-Turn Jailbreaking

11/1/2025

Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models

Affects: claude-3-haiku, gemini-1.5-flash, gemini-1.5-pro, gemini-2.0-flash, gpt-4o-mini, gpt-3.5-turbo, llama2-7b, llama2-13b, llama3-8b, deepseek-chat, mistral-7b-instruct-v0.3, vicuna-13b-v1.5-16k

Personalized Disinformation Jailbreak Escalation

11/11/2025

Appending simple demographic persona details to prompts requesting policy-violating content can bypass the safety mechanisms of Large Language Models. This technique, referred to as persona-targeted prompting, adds details such as country, generation, and political orientation to a request for a harmful narrative (e.g., disinformation). This systematically increases the jailbreak rate across most tested models and languages, in some cases by over 10 percentage points, enabling the generation of harmful content that would otherwise be refused.

A Multilingual, Large-Scale Study of the Interplay between LLM Safeguards, Personalisation, and Disinformation

Affects: gpt-40, claude-3.5-sonnet, grok-2, llama-3-8b-instruct, gemma-2-9b-instruct, mistral-nemo-instruct, qwen-2.5-7b-instruct, vicuna-1.5-7b-instruct

Persuasive Jailbreak Fingerprint

11/1/2025

Large Language Models (LLMs) are vulnerable to jailbreak attacks that use persuasive techniques grounded in social psychology to bypass safety alignments. Malicious instructions can be reframed using one of Cialdini's seven principles of persuasion (Authority, Reciprocity, Commitment, Social Proof, Liking, Scarcity, and Unity). These rephrased prompts, which remain human-readable and can be generated automatically, manipulate the LLM into complying with harmful requests it would otherwise refuse. The attack's effectiveness varies by principle and by model, revealing distinct "persuasive fingerprints" of susceptibility.

Uncovering the Persuasive Fingerprint of LLMs in Jailbreaking Attacks

Affects: wizardlm-uncensored2, vicuna, llama2, llama3, gemma3, deepseek-r1, phi-4, gpt-2

Schema Exploitation Jailbreak

10/31/2025

A vulnerability exists in Large Language Models where their strong adherence to processing structured data schemas can be exploited to bypass safety mechanisms. The attack, named BreakFun, uses a multi-component prompt that combines an innocent framing, a Chain-of-Thought (CoT) instruction, and a core "Trojan Schema." This schema is an adversarially designed data structure (e.g., a Python class definition) that embeds a harmful user request. By instructing the model to simulate the hypothetical output of code that uses this schema, the model's cognitive resources are misdirected towards fulfilling the structural and syntactic requirements of the task, causing it to overlook and comply with the embedded harmful request.

BreakFun: Jailbreaking LLMs via Schema Exploitation

Affects: gpt-4.1 mini, gemini 2.5 flash, claude-3.5 sonnet, claude-3 haiku, kimi-k2, ernie-4.5, gpt-oss 20b, deepseek-r1, gemma3 12b, qwen3 8b, llama 3.1 8b, mistral 7b, zephyr 7b, qwen3-max

Special Token Jailbreak

10/31/2025

Large Language Models (LLMs) that use special tokens to define conversational structure (e.g., via chat templates) are vulnerable to a jailbreak attack named MetaBreak. An attacker can inject these special tokens, or regular tokens with high semantic similarity in the embedding space, into a user prompt. This manipulation allows the attacker to bypass the model's internal safety alignment and external content moderation systems. The attack leverages four primitives:

Response Injection: Forging an assistant's turn within the user prompt to trick the model into believing it has already started to provide an affirmative response.
Turn Masking: Using a few-shot, word-by-word construction to make the injected response resilient to disruption from platform-added chat template wrappers.
Input Segmentation: Splitting sensitive keywords with injected special tokens to evade detection by content moderators, which may fail to reconstruct the original term, while the more capable target LLM can.
Semantic Mimicry: Bypassing special token sanitization defenses by substituting them with regular tokens that have a minimal L2 norm distance in the embedding space, thereby preserving the token's structural function.

MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation

Affects: llama-3.3-70b-instruct, qwen-2.5-72b-instruct, gemma-2-27b-instruct, phi4-14b, llamaguard, llama-3.1-405b, llama-3.1-8b, gpt-4.1, claude-opus-4, llamaguard3-8b, promptguard-86m, shieldgemma2-27b

Page 1 of 36