Published 6/1/2023
Analyzed 12/29/2024
A vulnerability in vision-integrated Large Language Models (VLMs) allows an attacker to circumvent safety mechanisms through the use of adversarially crafted visual examples. A single, carefully constructed image can universally "jailbreak" the model, causing it to generate harmful content in response to a wide range of subsequent prompts, even those not included in the adversarial example's training data. This vulnerability extends beyond simple misclassification to encompass the execution of…
Visual adversarial examples jailbreak large language models
Evaluated models: InstructBLIP, MiniGPT-4
Source: arXiv