Cross-Modal Entanglement Jailbreak
A vulnerability in advanced Vision-Language Models (VLMs) allows attackers to bypass safety alignment mechanisms via a Cross-Modal Entanglement Attack (COMET). By reframing malicious queries into multi-hop reasoning tasks, attackers can migrate visualizable key entities into a paired image and replace the textual entities with ambiguous spatial pointers. This forces the VLM to reconstruct the harmful intent through its own self-induced cross-modal reasoning, effectively bypassing filters that…
Evaluated models: GPT-4.1, GPT-4.1 Mini, Gemini 2.5 Flash+6 more
- GPT-4.1
- GPT-4.1 Mini
- Gemini 2.5 Flash
- Qwen3-VL 235B-A22B Instruct
- Qwen 2.5 VL 72B Instruct
- GLM-4.5V
- Gemini 2.5 Pro
- Llama 4 Maverick
- Claude Haiku 4.5