Skip to main content
LLM Security Database
Skip to research search
Last analyzed 9/9/2026

Language Model Security Database

985 research findings · 1123 evaluated models

Filtered research findings

267 entries

Matches every word across titles, descriptions, sources, affected systems, and models.

Published 10/1/2024
Analyzed 12/29/2024

Refusal-trained Large Language Models (LLMs) show decreased safety when deployed as browser agents compared to their performance in chatbot settings. Attack methods effective at jailbreaking LLMs in chat contexts also successfully bypass safety mechanisms in browser agents, leading to the execution of harmful behaviors. This vulnerability stems from a lack of generalization of safety training to agentic, real-world interaction scenarios and the increased context available to the agent (browser…

Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents
Evaluated models: o1-preview, o1-mini, GPT-4 Turbo +5 more

Source: arXiv

Published 10/1/2024
Analyzed 12/29/2024

Large Language Models (LLMs) are vulnerable to transferable ensemble black-box jailbreak attacks. The vulnerability allows an attacker to bypass safety mechanisms and elicit undesired or harmful responses from the LLM by using an ensemble of LLM-as-attacker methods that optimize malicious prompts, adaptively adjusting resources based on prompt difficulty, and strategically modifying prompt semantics to evade detection.

Transferable Ensemble Black-box Jailbreak Attacks on Large Language Models
Evaluated models: Deepseek-v2.5, Gemma 2B IT, Gemma 2 9B IT +5 more

Source: arXiv

Published 10/1/2024
Analyzed 12/28/2024

Large language models (LLMs) controlling robots are vulnerable to jailbreaking attacks. The ROBOPAIR algorithm demonstrates that malicious prompts can bypass safety mechanisms, causing robots to perform harmful physical actions. This vulnerability exploits the LLM's reliance on textual prompts and its potential lack of sufficient contextual understanding to prevent unsafe commands. The attack is effective across different access levels.

Jailbreaking LLM-controlled robots
Evaluated models: GPT-3.5 Turbo, GPT-4, GPT-4o +1 more

Source: arXiv

Published 9/1/2024
Analyzed 12/29/2024

Jailbreaking vulnerabilities in Large Language Models (LLMs) used in Retrieval-Augmented Generation (RAG) systems allow escalation of attacks from entity extraction to full document extraction and enable the propagation of self-replicating malicious prompts ("worms") within interconnected RAG applications. Exploitation leverages prompt injection to force the LLM to return retrieved documents or execute malicious actions specified within the prompt.

Unleashing worms and extracting data: Escalating the outcome of attacks against rag-based inference in scale and severity using jailbreaking
Evaluated models: Gemini 1.5 Flash

Source: arXiv

Published 9/1/2024
Analyzed 12/29/2024

Large Language Models (LLMs) used in role-playing systems are vulnerable to character hallucination attacks, a form of jailbreak exploiting "query sparsity" and "role-query conflict". Query sparsity occurs when prompts fall outside the model's training data distribution, causing it to generate out-of-character responses. Role-query conflict arises when the prompt contradicts the established character persona, leading to inconsistent behavior. These vulnerabilities allow attackers to elicit…

RoleBreak: Character Hallucination as a Jailbreak Attack in Role-Playing Systems
Evaluated models: Claude 3 Haiku, GPT-3.5 Turbo, Llama 3 8B +1 more

Source: arXiv

Published 8/1/2024
Analyzed 12/29/2024

The ALERT-Motion framework demonstrates a vulnerability in text-to-motion (T2M) models where an attacker can craft subtly modified text prompts (adversarial prompts) that cause the model to generate motions significantly different from those intended by the benign prompt, yet semantically similar to a target motion specified by the attacker. The attack leverages a large language model (LLM) to autonomously generate these adversarial prompts, bypassing simple keyword-based detection mechanisms…

Autonomous LLM-Enhanced Adversarial Attack for Text-to-Motion
Evaluated models: Mdm, Mld

Source: arXiv

Published 8/1/2024
Analyzed 12/28/2024

A vulnerability allows bypassing safety filters in text-to-image (T2I) models using a multi-agent framework ("Atlas") powered by Large Language Models (LLMs). Atlas iteratively generates and refines prompts, leveraging a Vision-Language Model (VLM) to assess filter activation and an LLM to select effective prompts that maintain semantic similarity to the original, malicious prompt while evading the filter. This enables the generation of images containing unsafe content.

Jailbreaking text-to-image models with llm-based agents
Evaluated models: DALL-E 3, LLaVA 1.5 13B, Sharegpt4v-13B +4 more

Source: arXiv

Published 7/1/2024
Analyzed 12/29/2024

LLM-based autonomous agents are vulnerable to malfunction amplification attacks. These attacks exploit the inherent instability of agents by inducing repetitive or irrelevant actions through various methods including prompt injection and adversarial perturbations, leading to agent malfunction and task failure. The attacks do not rely on overtly harmful actions, making them harder to detect with standard LLM safety mechanisms.

Breaking agents: Compromising autonomous llm agents through malfunction amplification
Evaluated models: Claude 2, GPT-3.5 Turbo, GPT-4

Source: arXiv

Published 7/1/2024
Analyzed 12/29/2024

Embodied Large Language Models (LLMs) are vulnerable to manipulation via voice-based interactions, leading to the execution of harmful physical actions. Attacks exploit three vulnerabilities: (1) cascading LLM jailbreaks resulting in malicious robotic commands; (2) misalignment between linguistic outputs (verbal refusal) and physical actions (command execution); and (3) conceptual deception, where seemingly benign instructions lead to harmful outcomes due to incomplete world knowledge within…

BadRobot: Manipulating Embodied LLMs in the Physical World
Evaluated models: BERT, GPT-3.5 Turbo, GPT-4 Turbo +2 more

Source: arXiv

Published 7/1/2024
Analyzed 12/28/2024

A vulnerability in Retrieval-Augmented Generation (RAG)-based Large Language Model (LLM) agents allows attackers to inject malicious demonstrations into the agent's memory or knowledge base. By crafting a carefully optimized trigger, an attacker can manipulate the agent's retrieval mechanism to preferentially retrieve these poisoned demonstrations, causing the agent to produce adversarial outputs or take malicious actions even when seemingly benign prompts are used. The attack, termed…

Agentpoison: Red-teaming llm agents via poisoning memory or knowledge bases
Evaluated models: GPT-2, GPT-3.5 Turbo, Llama 3 70B +2 more

Source: arXiv

Research methodology

Entries summarize publicly available primary-source security research. Model names reflect only systems explicitly evaluated by the cited paper, and measurements are research-reported unless independent verification is stated.