Skip to main content
LLM Security Database
Skip to research search
Last analyzed 9/9/2026

Language Model Security Database

985 research findings · 1123 evaluated models

Latest research findings

985 entries

Matches every word across titles, descriptions, sources, affected systems, and models.

Controlled safety-tuning experiments link boilerplate refusal statements to unnecessary refusals of benign requests. Request-specific rationales improve benign compliance, with benchmark-dependent safety tradeoffs.

Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models
Evaluated models: Llama 3.1 8B, Mistral 7B v0.3, Gemma 2 9B +9 more

Source: arXiv

Deleting a memory record can leave its information in an agent's summaries, pending plans and KV cache. The paper evaluates revocation across this derived execution state.

Forgetting Without Restarting: Execution-State Unlearning for Stateful LLM Agents
Evaluated models: Llama 3.1 8B Instruct, Qwen 2.5 7B Instruct, Mistral 7B Instruct v0.3

Source: arXiv

Published 9/4/2026
Analyzed 9/9/2026

KoNA measures whether vision-language models answer valid image questions while refusing unsafe components or correcting unsupported premises. Its 9,300 question-answer pairs include mixed and fully answerable controls.

Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models
Evaluated models: InternVL3 2B Instruct, InternVL3-78B-Instruct, Qwen 2.5 VL 3B Instruct +5 more

Source: arXiv

Published 9/2/2026
Analyzed 9/9/2026

Document-image PII detection can miss identifiers despite improving average localization scores. LeakageBench evaluates 500 pages with 11,954 annotations.

LeakageBench: Document-Level Leakage Risk for Redacting Personally Identifiable Information in Document Images
Evaluated models: GLiNER-base, GLiNER-multi-PII, NVIDIA GLiNER PII +5 more

Source: arXiv

Published 8/26/2026
Analyzed 9/9/2026

A compromised serving framework can violate user-data isolation through shared GPU state. GIFT evaluates per-user information-flow enforcement in vLLM and DistServe.

Here is a GIFT: Enforcing User Data Isolation in LLM Serving via GPU Information Flow Tracking
Evaluated models: Qwen 2.5 14B, Qwen 2.5 32B, Qwen 2.5 72B +3 more

Source: arXiv

Published 8/24/2026
Analyzed 9/9/2026

A stored interaction can later steer a memory-augmented agent's answer without direct memory-store access. The study evaluates persistent response manipulation in MemoryOS and MemGPT.

InjecMEM: Memory Injection Attack on LLM Agent Memory Systems
Evaluated models: Qwen 2.5 7B Instruct, Qwen 2.5 1.5B Instruct, Qwen 2.5 3B Instruct +3 more

Source: arXiv

Research methodology

Entries summarize publicly available primary-source security research. Model names reflect only systems explicitly evaluated by the cited paper, and measurements are research-reported unless independent verification is stated.