Multimodal Risk Diffusion Jailbreak
Multimodal Large Language Models (MLLMs) are vulnerable to a heuristic-induced multimodal risk distribution jailbreak attack. The attack successfully circumvents safety mechanisms by distributing malicious prompts across text and image modalities, preventing detection of harmful intent within either modality alone. An auxiliary LLM generates prompts to guide the target MLLM into reconstructing the malicious prompt and producing the desired harmful output.
Evaluated models: Deepseek-vl7B-chat, Gemini 1.5 Pro, Glm-4v-9B+7 more
- Deepseek-vl7B-chat
- Gemini 1.5 Pro
- Glm-4v-9B
- GPT-4o-0513
- Llava v1.5-7B
- Llava v1.6-mistral-7B-hf
- MiniGPT-4
- Qwen VL Chat
- Qwen VL Max
- Yi-vl-34B