Benign Audio Jailbreak
Audio-Language Models (ALMs) including Qwen2.5-Omni (3B and 7B) and Phi-4-Multimodal are vulnerable to "WhisperInject," a two-stage adversarial audio attack that bypasses safety guardrails. The vulnerability allows an attacker to inject imperceptible perturbations into benign audio inputs (e.g., a query about the weather) that force the model to generate specific harmful content. The attack utilizes a novel optimization method, Reinforcement Learning with Projected Gradient Descent (RL-PGD)…
Evaluated models: Qwen 2.5 Omni 3B, Qwen 2.5 Omni 7B, Phi-4 Multimodal+2 more
- Qwen 2.5 Omni 3B
- Qwen 2.5 Omni 7B
- Phi-4 Multimodal
- Gemma 3n E2B
- Gemma 3n E4B