Controlled safety-tuning experiments link boilerplate refusal statements to unnecessary refusals of benign requests. Request-specific rationales improve benign compliance, with benchmark-dependent safety tradeoffs.
Source: arXiv
Explore primary-source AI security research, evaluated models, attack techniques, and defensive evidence.
985 research findings · 1123 evaluated models
Matches every word across titles, descriptions, sources, affected systems, and models.
Controlled safety-tuning experiments link boilerplate refusal statements to unnecessary refusals of benign requests. Request-specific rationales improve benign compliance, with benchmark-dependent safety tradeoffs.
Source: arXiv
Deleting a memory record can leave its information in an agent's summaries, pending plans and KV cache. The paper evaluates revocation across this derived execution state.
Source: arXiv
Exposed retrieval embeddings can reveal source text. SHAQ evaluates indexing generated queries instead of direct document embeddings to reduce that leakage.
Source: arXiv
KoNA measures whether vision-language models answer valid image questions while refusing unsafe components or correcting unsupported premises. Its 9,300 question-answer pairs include mixed and fully answerable controls.
Source: arXiv
Untrusted document images can redirect vision-language agents across instruction and tool-authorization boundaries. Repeat-After-Me evaluates six victim models on constructed document tasks.
Source: arXiv
Document-image PII detection can miss identifiers despite improving average localization scores. LeakageBench evaluates 500 pages with 11,954 annotations.
Source: arXiv
Context assembly can promote repository, tool or skill content into higher-priority instructions or persistent state. The paper studies 12 pinned agent-harness versions.
Source: arXiv
A compromised serving framework can violate user-data isolation through shared GPU state. GIFT evaluates per-user information-flow enforcement in vLLM and DistServe.
Source: arXiv
A stored interaction can later steer a memory-augmented agent's answer without direct memory-store access. The study evaluates persistent response manipulation in MemoryOS and MemGPT.
Source: arXiv
Third-party skills can inflate coding-agent resource use while an otherwise legitimate task remains functional.
Source: arXiv
Entries summarize publicly available primary-source security research. Model names reflect only systems explicitly evaluated by the cited paper, and measurements are research-reported unless independent verification is stated.