MLLM Over-Reasoning Safety Risk
Multimodal Large Language Models (MLLMs) exhibit a vulnerability to "Reasoning-based Multi-Image Attacks," where safety guardrails are bypassed by distributing harmful intent across multiple images (2–4 inputs). Unlike single-image jailbreaks that rely on visual obfuscation, this vulnerability exploits the model's reasoning capabilities. By presenting images that share a specific relationship (e.g., Temporal Jump, Spatial Juxtaposition, or Causality), an attacker can compel the model to infer…
Evaluated models: GPT-4o, GPT-4o Mini, Gemini 1.5 Pro+11 more
- GPT-4o
- GPT-4o Mini
- Gemini 1.5 Pro
- Gemini 1.5 Flash
- Qwen 2.5 VL 3B Instruct
- Qwen 2.5 VL 32B Instruct
- LLaVA 1.5 7B
- Llama 3 LLaVA-NeXT 8B
- InternVL3 8B
- InternVL3 38B
- InternVL3 78B
- MiniCPM-o 2.6
- Skywork-R1V3 38B
- GLM-4.1V 9B Thinking