Skip to main content
LLM Security Database
Skip to research search
Last analyzed Sep 9, 2026

Language Model Security Database

985 research findings · 1119 evaluated models

Research findings

985 entries

Matches every word across titles, descriptions, sources, affected systems, and models.

Boilerplate safety tuning causes false refusals on benign requests

Published Analyzed Source: arXiv

Controlled safety-tuning experiments link boilerplate refusal statements to unnecessary refusals of benign requests. Request-specific rationales improve benign compliance, with benchmark-dependent safety tradeoffs.

Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models
Evaluated models: Llama 3.1 8B, Mistral 7B v0.3, Gemma 2 9B+9 more
  • Llama 3.1 8B
  • Mistral 7B v0.3
  • Gemma 2 9B
  • Qwen 2.5 7B
  • Gemma 2 2B
  • Qwen 2.5 3B
  • Llama 3.1 70B
  • Qwen 2.5 72B
  • Llama 3.1 8B Instruct
  • Mistral 7B Instruct v0.3
  • Gemma 2 9B IT
  • Qwen 2.5 7B Instruct

Deleted agent memory persists in live execution state

Published Analyzed Source: arXiv

Deleting a memory record can leave its information in an agent's summaries, pending plans and KV cache. The paper evaluates revocation across this derived execution state.

Forgetting Without Restarting: Execution-State Unlearning for Stateful LLM Agents
Evaluated models: Llama 3.1 8B Instruct, Qwen 2.5 7B Instruct, Mistral 7B Instruct v0.3

Selective refusal gaps in visual question answering

Published Analyzed Source: arXiv

KoNA measures whether vision-language models answer valid image questions while refusing unsafe components or correcting unsupported premises. Its 9,300 question-answer pairs include mixed and fully answerable controls.

Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models
Evaluated models: InternVL3 2B Instruct, InternVL3-78B-Instruct, Qwen 2.5 VL 3B Instruct+5 more
  • InternVL3 2B Instruct
  • InternVL3-78B-Instruct
  • Qwen 2.5 VL 3B Instruct
  • Qwen 2.5 VL 72B Instruct
  • GPT-5
  • Gemini 2.5 Flash
  • InternVL3-2B-KoNA
  • Qwen2.5-VL-3B-KoNA

Visual prompt injection crosses document trust boundaries

Published Analyzed Source: arXiv

Untrusted document images can redirect vision-language agents across instruction and tool-authorization boundaries. Repeat-After-Me evaluates six victim models on constructed document tasks.

Repeat-After-Me: Black-Box Adaptive Visual Prompt Injection
Evaluated models: claude-opus-4-7, GPT-5.5, Gemini 3.1 Pro+3 more
  • claude-opus-4-7
  • GPT-5.5
  • Gemini 3.1 Pro
  • Qwen3.6-27B
  • Qwen3-VL 32B Instruct
  • InternVL3.5 38B Instruct

PII leakage in document-image redaction benchmarks

Published Analyzed Source: arXiv

Document-image PII detection can miss identifiers despite improving average localization scores. LeakageBench evaluates 500 pages with 11,954 annotations.

LeakageBench: Document-Level Leakage Risk for Redacting Personally Identifiable Information in Document Images
Evaluated models: GLiNER-base, GLiNER-multi-PII, NVIDIA GLiNER PII+5 more
  • GLiNER-base
  • GLiNER-multi-PII
  • NVIDIA GLiNER PII
  • GLiNER2
  • Qwen3-VL-32B
  • InternVL3 38B
  • GPT-5.4
  • GPT-5.5

Context privilege escalation in AI agent harnesses

Published Analyzed Source: arXiv

Context assembly can promote repository, tool or skill content into higher-priority instructions or persistent state. The paper studies 12 pinned agent-harness versions.

What's in Your Agent's Context? Context Privilege Escalation Attacks against AI Agent Harness
Evaluated models: GPT-5.5, GPT-5.4-mini, Claude Sonnet 4.6+4 more
  • GPT-5.5
  • GPT-5.4-mini
  • Claude Sonnet 4.6
  • Claude Opus 4.6
  • Gemini 2.5 Flash
  • Gemini 2.5 Pro
  • DeepSeek V4 Flash

Cross-user data isolation in shared GPU serving

Published Analyzed Source: arXiv

A compromised serving framework can violate user-data isolation through shared GPU state. GIFT evaluates per-user information-flow enforcement in vLLM and DistServe.

Here is a GIFT: Enforcing User Data Isolation in LLM Serving via GPU Information Flow Tracking
Evaluated models: Qwen 2.5 14B, Qwen 2.5 32B, Qwen 2.5 72B+3 more
  • Qwen 2.5 14B
  • Qwen 2.5 32B
  • Qwen 2.5 72B
  • OPT 13B
  • OPT 30B
  • OPT 66B

Single-interaction poisoning of agent memory

Published Analyzed Source: arXiv

A stored interaction can later steer a memory-augmented agent's answer without direct memory-store access. The study evaluates persistent response manipulation in MemoryOS and MemGPT.

InjecMEM: Memory Injection Attack on LLM Agent Memory Systems
Evaluated models: Qwen 2.5 7B Instruct, Qwen 2.5 1.5B Instruct, Qwen 2.5 3B Instruct+3 more
  • Qwen 2.5 7B Instruct
  • Qwen 2.5 1.5B Instruct
  • Qwen 2.5 3B Instruct
  • Qwen 2.5 14B Instruct
  • Llama 3.1 8B Instruct
  • Mistral 7B Instruct v0.3

Research methodology

Entries summarize publicly available primary-source security research. Model names reflect only systems explicitly evaluated by the cited paper, and measurements are research-reported unless independent verification is stated.

Discovery and publication are separate: new papers enter a candidate queue, then require human review and explicit approval before publication. Updates follow that review process, not a guaranteed daily publishing schedule. “Last analyzed” is the newest analysis date among published entries; it does not mean that all recent research has been covered.