Benign LLM Secondary Risks
Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) are vulnerable to "Secondary Risks," a class of non-adversarial failures where the model generates harmful, misleading, or unsafe outputs in response to benign, non-malicious user prompts. Unlike jailbreaks which require adversarial inputs, secondary risks arise from imperfect generalization and alignment failures during standard interactions. This vulnerability manifests primarily in two primitives: 1. Excessive…
Evaluated models: GPT-4o, Claude 3.7 Sonnet, GPT-4 Turbo+9 more
- GPT-4o
- Claude 3.7 Sonnet
- GPT-4 Turbo
- Gemini 2.0 Pro
- DeepSeek V3
- Llama 3.3 70B
- Qwen 2.5 32B
- Phi-4
- Gemma 2 27B IT
- LLaVA OneVision Qwen2 72B
- Pixtral 12B
- MiniCPM-o 2.6