Skip to main content
LLM Security Database
Skip to research search
Last analyzed 9/9/2026

Language Model Security Database

985 research findings · 1123 evaluated models

Filtered research findings

736 entries

Matches every word across titles, descriptions, sources, affected systems, and models.

Published 11/1/2025
Analyzed 12/8/2025

Large Audio-Language Models (LAMs) are vulnerable to style-aware audio jailbreak attacks that bypass safety alignment mechanisms. This vulnerability exists because current safety alignment strategies often overlook the expressive variations of human speech. Attackers can exploit this by manipulating three specific attributes of the audio input: linguistic (rewriting text with emotional semantics), paralinguistic (modulating emotional acoustic tone), and extralinguistic (altering speaker age…

StyleBreak: Revealing Alignment Vulnerabilities in Large Audio-Language Models via Style-Aware Audio Jailbreak
Evaluated models: GPT-4o, Llama 3.1 8B, Qwen 2 7B +1 more

Source: arXiv

Published 11/1/2025
Analyzed 12/1/2025

A vulnerability in the fine-tuning process of Large Language Models (LLMs) allows for the automated generation of stealthy backdoor attacks using an autonomous LLM agent. This method, termed AutoBackdoor, creates a pipeline to generate semantically coherent trigger phrases and corresponding poisoned instruction-response pairs. Unlike traditional backdoor attacks that rely on fixed, often anomalous triggers, this technique produces natural language triggers that are contextually relevant and…

AutoBackdoor: Automating Backdoor Attacks via LLM Agents
Evaluated models: GPT-4o, GPT-4o Mini, Llama 3.1 8B Instruct +3 more

Source: arXiv

Published 11/1/2025
Analyzed 12/8/2025

Improper restriction of the "Capability Space" in Large Language Model (LLM) applications allows remote attackers to manipulate application behavior through "Goal Deviation" attacks. This vulnerability arises when developers rely on the broad capabilities of a foundational model (e.g., GPT-4, LLaMA) without implementing sufficient negative constraints or disabling default plugins (e.g., DALL-E, Web Search) in the system prompt. Attackers can exploit this via natural language inputs to trigger…

Beyond Jailbreak: Unveiling Risks in LLM Applications Arising from Blurred Capability Boundaries
Evaluated models: Not reported

Source: arXiv

Published 11/1/2025
Analyzed 12/8/2025

Large Language Models (LLMs) from multiple vendors exhibit vulnerabilities to jailbreaking techniques that bypass safety guardrails, enabling the automated generation of highly persuasive phishing content specifically targeted at elderly victims. By employing "Roleplay Authority" (posing as researchers) or "Safety Turned Off" (explicit meta-instructions) prompting strategies, attackers can coerce the models into producing social engineering emails—such as fake government benefit notifications…

Can AI Models be Jailbroken to Phish Elderly Victims? An End-to-End Evaluation
Evaluated models: GPT-5, Claude Sonnet 4, Gemini 2.5 Pro +3 more

Source: arXiv

Published 11/1/2025
Analyzed 12/9/2025

Large Language Models (LLMs), specifically GPT-4o, GPT-4o-mini, LLaMA-2-13B, Mistral-7B, and Phi-3.5-mini, are vulnerable to Man-in-the-Middle (MitM) adversarial prompt injections that undermine factual recall. Termed the "$\chi$mera" (Chimera) attack framework, this vulnerability exists when an attacker intercepts and modifies user queries (e.g., via malicious browser extensions, compromised frontends, or proxy middleware) before they reach the victim model. By appending adversarial…

Injecting Falsehoods: Adversarial Man-in-the-Middle Attacks Undermining Factual Recall in LLMs
Evaluated models: GPT-4o, Llama 2 13B, Mistral 7B +1 more

Source: arXiv

Published 11/1/2025
Analyzed 12/8/2025

Large Language Models (LLMs), specifically GPT-3.5-turbo, LLaMA3-8B-instruct, and DeepSeek-R1-Distill-Qwen-7B, are vulnerable to a "Self-Harm" jailbreak attack (Self-HarmLLM). This vulnerability exploits the model's ability to understand its own safety boundaries to generate adversarial inputs against itself. An attacker utilizes a two-session approach: in the first session (Mitigation Session), the attacker instructs the model to rewrite a harmful query into a "Mitigated Harmful Query"…

Self-HarmLLM: Can Large Language Model Harm Itself?
Evaluated models: GPT-3.5 Turbo, Llama 3 8B Instruct, DeepSeek R1 Distill Qwen 7B

Source: arXiv

Published 11/1/2025
Analyzed 12/5/2025

A vulnerability exists in certain Large Language Models and diffusion models due to discontinuities in their latent space, which arise from data sparsity during training. An attacker can craft inputs containing lexically rare or semantically ambiguous constructs to guide the model's inference process toward these unstable, poorly-conditioned regions. This technique, termed "Alignment Degradation Induction," can degrade or bypass safety alignment mechanisms. Through iterative, multi-turn…

Exploiting Latent Space Discontinuities for Building Universal LLM Jailbreaks and Data Extraction Attacks
Evaluated models: Not reported

Source: arXiv

Published 11/1/2025
Analyzed 12/9/2025

Large Language Models (LLMs) are vulnerable to Linguistic Style Jailbreaks, a technique where an attacker reframes a harmful prompt using specific linguistic tones—such as politeness, fear, curiosity, or compassion—to bypass safety guardrails. While standard safety alignment (RLHF) effectively filters harmful requests phrased in neutral or hostile tones, it fails to generalize to prompts where the semantic intent remains harmful but the stylistic framing triggers compliant, helpful, or…

Say It Differently: Linguistic Styles as Jailbreak Vectors
Evaluated models: Llama 3.1 8B Instruct, Llama 3.2 1B Instruct, Llama 3.2 3B Instruct +13 more

Source: arXiv

Published 11/1/2025
Analyzed 11/20/2025

A vulnerability in Large Language Models (LLMs) allows for systematic jailbreaking through a meta-optimization framework called AMIS (Align to MISalign). The attack uses a bi-level optimization process to co-evolve both the jailbreak prompts and the scoring templates used to evaluate them.

Align to Misalign: Automatic LLM Jailbreak with Meta-Optimized LLM Judges
Evaluated models: Claude 3.5 Haiku, Claude 3.5 Sonnet, Claude Sonnet 4 +3 more

Source: arXiv

Published 11/1/2025
Analyzed 12/9/2025

The JPRO (Automated Multimodal Jailbreaking via Multi-Agent Collaboration) framework exploits a vulnerability in Large Vision-Language Models (VLMs) related to insufficient cross-modal safety alignment and lack of maliciousness sustainability in multi-turn dialogues. The attack leverages a multi-agent system (Planner, Attacker, Modifier, Verifier) to automate the generation of adversarial image-text pairs. By employing hybrid tactics—such as combining role-playing with malicious content…

JPRO: Automated Multimodal Jailbreaking via Multi-Agent Collaboration Framework
Evaluated models: GPT-4o, GPT-4o Mini, GPT-4.1 +3 more

Source: arXiv

Research methodology

Entries summarize publicly available primary-source security research. Model names reflect only systems explicitly evaluated by the cited paper, and measurements are research-reported unless independent verification is stated.