The LMVD-ID is an internal research identifier, not an official CVE identifier.
Iterative Semantic Jailbreak
The MIST attack exploits a vulnerability in black-box large language models (LLMs) allowing iterative semantic tuning of prompts to elicit harmful responses. The attack leverages synonym substitution and optimization…
Paper-evaluated models(6)
Claude 3.5 Sonnet, GPT-4 Turbo, GPT-4o +3 more
- Claude 3.5 Sonnet
- GPT-4 Turbo
- GPT-4o
- GPT-4o Mini
- Llama 2 7B Chat
- Vicuna 7B v1.5
Description
The MIST attack exploits a vulnerability in black-box large language models (LLMs) allowing iterative semantic tuning of prompts to elicit harmful responses. The attack leverages synonym substitution and optimization strategies to bypass safety mechanisms without requiring access to the model's internal parameters or weights. The vulnerability lies in the susceptibility of the LLM to semantically similar prompts that trigger unsafe outputs.
Examples
See arXiv:2506.16792v1. The paper includes examples demonstrating the iterative semantic tuning process and successful jailbreaks across multiple open-source and closed-source LLMs.
Impact
Successful exploitation of this vulnerability allows attackers to circumvent built-in safety mechanisms and obtain harmful outputs from LLMs, including generation of illegal instructions, toxic content, or personally identifiable data leakage. The attack's ability to transfer to various models further magnifies the impact. The low query overhead indicates practical feasibility.
Affected Systems
The vulnerability affects a wide range of LLMs, including (but not limited to) Vicuna-7B-v1.5, Llama-2-7B-chat, Claude-3.5-sonnet, GPT-4o-mini, GPT-4o-0806, and GPT-4-turbo. The attack's transferability suggests that many other LLMs are potentially vulnerable.
Mitigation Steps
- Implement robust semantic similarity checks to detect slight variations in prompts designed to elicit harmful outputs.
- Develop and incorporate defenses that are resilient to synonym substitution and variations in prompt phrasing.
- Integrate advanced detection mechanisms using gradient analyses or other methods that analyze prompt and output relationships, going beyond simple keyword filtering.
- Regularly update safety filters and mechanisms based on evolving attack techniques.
- Conduct comprehensive adversarial testing and incorporate findings into model development and training.
Research context and confidence
- Evidence and verification
- Paper-reported; independent reproduction is not documented.
- Primary research source linked.
- Severity
- Not rated by this catalog.
- Source and publication type
- arXiv · Research preprint.
- Peer-review status is not provided by this source.
- Author and publication status
- Author metadata is not stored; see the primary paper.
- Threat model and attacker access
- Black-box model, service, or application access.
- Related deployment categories
- No related deployment category is classified.
- Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
- Affected systems
- The vulnerability affects a wide range of LLMs, including (but not limited to) Vicuna-7B-v1.5, Llama-2-7B-chat, Claude-3.5-sonnet, GPT-4o-mini, GPT-4o-0806, and GPT-4-turbo. The attack's transferability suggests that…
Research Paper
MIST: Jailbreaking Black-box Large Language Models via Iterative Semantic Tuning
Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.
View PaperEvidence
This entry is based on a primary research source. Its findings are paper-reported; independent reproduction and verification are not claimed.
https://arxiv.org/abs/2506.16792Related research
- Academic Paper Trust Jailbreak
Published July 1, 2025 · model-layer, prompt-layer, injection
- Agentic Red-Teaming Uncovers Novel Jailbreaks
Published June 1, 2025 · model-layer, prompt-layer, jailbreak
- Segmented Prompt Jailbreak
Published March 1, 2025 · prompt-layer, jailbreak, blackbox