Skip to main content
LLM Security Database
Skip to research search
Last analyzed 8/13/2026

Language Model Security Database

969 research findings · 1102 evaluated models

Filtered research findings

604 entries

Matches every word across titles, descriptions, sources, affected systems, and models.

Published 12/1/2023
Analyzed 12/29/2024

Large Language Models (LLMs) exhibit an inherent response tendency, predisposing them towards affirmation or rejection of instructions. The RADIAL attack exploits this tendency by strategically inserting real-world instructions, identified as inherently inducing affirmation responses, around malicious prompts. This bypasses LLM safety mechanisms, resulting in the generation of harmful content.

Analyzing the inherent response tendency of llms: Real-world instructions-driven jailbreak
Evaluated models: Baichuan 2 13B Chat, Baichuan 2 7B Chat, ChatGLM2 6B +3 more

Source: arXiv

Published 11/1/2023
Analyzed 12/28/2024

A vulnerability exists in large language models (LLMs) utilizing in-context learning (ICL). Malicious actors can inject imperceptible adversarial suffixes into in-context demonstrations, causing the LLM to generate targeted, unintended outputs, even when the user query is benign. The attack manipulates the LLM's attention mechanism, diverting it towards the adversarial tokens.

Hijacking large language models via adversarial in-context learning
Evaluated models: Llama 13B, Llama 3.1 8B, Llama 3.1 8B Instruct +3 more

Source: arXiv

Published 11/1/2023
Analyzed 12/29/2024

Large Language Models (LLMs) are vulnerable to jailbreaking attacks exploiting cognitive overload induced by multilingual prompts, veiled expressions, and effect-to-cause reasoning. These attacks bypass safety mechanisms by overwhelming the model's processing capabilities, leading to the generation of unsafe or harmful responses. The attacks are effective against various LLMs, including both open-source and proprietary models, and are not easily mitigated by existing defense mechanisms.

Cognitive overload: Jailbreaking large language models with overloaded logical thinking
Evaluated models: GPT-3.5 Turbo-0301, Guanaco 7B, Guanaco 13B +8 more

Source: arXiv

Published 11/1/2023
Analyzed 12/28/2024

A prompt injection vulnerability in OpenAI's custom GPT models allows attackers to extract the system prompt and potentially leak user-uploaded files. Attackers craft malicious prompts that manipulate the LLM into revealing sensitive information, even when defensive prompts are in place. The vulnerability is exacerbated when the model includes a code interpreter.

Assessing prompt injection risks in 200+ custom gpts
Evaluated models: Not reported

Source: arXiv

Published 11/1/2023
Analyzed 12/29/2024

A system prompt leakage vulnerability in GPT-4V allows extraction of internal system prompts through carefully crafted, incomplete conversations combined with image input. Extracted prompts can be used as highly effective jailbreak prompts, bypassing safety restrictions and leading to undesirable outputs, including revealing personally identifiable information from images.

Jailbreaking gpt-4v via self-adversarial attacks with system prompts
Evaluated models: GPT-4, GPT-4V, LLaVA 1.5

Source: arXiv

Published 11/1/2023
Analyzed 12/29/2024

Large Language Models (LLMs) are vulnerable to a targeted linguistic fuzzing attack that exploits the complexity of human language to bypass safety guardrails. The attack, termed "Jade," leverages transformational-generative grammar rules to systematically increase the syntactic complexity of benign seed questions, making them increasingly difficult for LLMs to recognize as malicious. This leads to the generation of unsafe content, even when the underlying semantics remain unchanged.

Jade: A linguistics-based safety evaluation platform for llm
Evaluated models: ChatGLM2 6B, GPT-2, GPT-3 +2 more

Source: arXiv

Published 11/1/2023
Analyzed 12/29/2024

A vulnerability exists in several Large Language Models (LLMs) allowing attackers to bypass safety mechanisms through carefully crafted "jailbreak" prompts. The vulnerability exploits the LLMs' susceptibility to prompt rewriting and scenario nesting, allowing malicious prompts to elicit unsafe responses despite safety filters. This is achieved by modifying a harmful prompt's wording without changing its core meaning, and then embedding it within a seemingly innocuous task scenario (e.g., code…

A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily
Evaluated models: Claude-instant-v1, Claude-v2, GPT-2 +4 more

Source: arXiv

Published 11/1/2023
Analyzed 1/26/2025

Large Language Models (LLMs) are vulnerable to a novel "DeepInception" attack that leverages the models' personification capabilities to bypass safety guardrails. The attack uses nested prompts to create a multi-layered fictional scenario, effectively hypnotizing the LLM into generating harmful content by exploiting its tendency towards obedience within the constructed narrative. This allows for continuous jailbreaks in subsequent interactions.

Deepinception: Hypnotize large language model to be jailbreaker
Evaluated models: GPT-3.5 Turbo, GPT-4, GPT-4o +1 more

Source: arXiv

Published 11/1/2023
Analyzed 12/29/2024

Large Language Models (LLMs) are vulnerable to persona modulation attacks, a black-box jailbreak technique that leverages an LLM assistant to generate prompts causing the target LLM to adopt harmful personas and produce unsafe outputs. This vulnerability circumvents built-in safety mechanisms, enabling the generation of responses related to illegal activities (e.g., synthesizing drugs, building bombs, money laundering), hate speech, and other harmful content. The attack's effectiveness is…

Scalable and transferable black-box jailbreaks for language models via persona modulation
Evaluated models: Claude 2, GPT-4

Source: arXiv

Published 11/1/2023
Analyzed 12/29/2024

Large Vision-Language Models (VLMs) are vulnerable to jailbreaking attacks via typographically rendered visual prompts. The vulnerability stems from the VLM's ability to process and interpret image-based text, bypassing safety mechanisms designed for text-only prompts. Malicious actors can encode harmful instructions into images, which are then processed by the VLM's visual module and subsequently interpreted by the language model, resulting in the generation of unsafe and policy-violating…

Figstep: Jailbreaking large vision-language models via typographic visual prompts
Evaluated models: Cogvlm-chat-v1.1, GPT-4V, Llava-v1.5-vicuna-v1.5-13B +4 more

Source: arXiv

Research methodology

Entries summarize publicly available primary-source security research. Model names reflect only systems explicitly evaluated by the cited paper, and measurements are research-reported unless independent verification is stated.