Skip to main content
LLM Security Database
Skip to research details
Back to research findings
LMVD-ID: 2774f631
Paper published November 1, 2023
Entry analyzed December 29, 2024
Paper-reported evidence
Confidence: Source-linked

The LMVD-ID is an internal research identifier, not an official CVE identifier.

Persona-Based LLM Jailbreak

Large Language Models (LLMs) are vulnerable to persona modulation attacks, a black-box jailbreak technique that leverages an LLM assistant to generate prompts causing the target LLM to adopt harmful personas and…

BibTeX citation

Paper-evaluated models(2)

  • Claude 2
  • GPT-4

Description

Large Language Models (LLMs) are vulnerable to persona modulation attacks, a black-box jailbreak technique that leverages an LLM assistant to generate prompts causing the target LLM to adopt harmful personas and produce unsafe outputs. This vulnerability circumvents built-in safety mechanisms, enabling the generation of responses related to illegal activities (e.g., synthesizing drugs, building bombs, money laundering), hate speech, and other harmful content. The attack's effectiveness is amplified by the assistant LLM's capabilities; more powerful assistants generate more effective jailbreaks.

Examples

See the paper for examples of harmful outputs generated using this technique against GPT-4, Claude 2, and Vicuna. Specific examples include detailed instructions for illegal activities and hate speech generation.

Impact

Successful exploitation allows attackers to bypass LLM safety measures and elicit harmful and illegal information, instructions, or content. The impact includes the potential for physical harm, financial loss, and the spread of misinformation and hate speech. The scalability of the attack, enabled by the use of an LLM assistant, significantly increases the threat.

Affected Systems

Large Language Models (LLMs) such as GPT-4, Claude 2, and Vicuna, and potentially other LLMs equipped with similar safety mechanisms are affected. The vulnerability is independent of the specific model architecture or training data.

Mitigation Steps

  • Improve detection of persona modulation attempts within the LLM's safety mechanisms.
  • Develop and implement more robust safety measures that are resistant to persona manipulation and prompt engineering techniques.
  • Limit or monitor the capabilities of LLM Assistants used in production settings to reduce the potential for malicious exploitation.
  • Implement mechanisms to detect and mitigate malicious prompt engineering using techniques like those described in this paper.

Research context and confidence

Evidence and verification
Paper-reported; independent reproduction is not documented.
Primary research source linked.
Severity
Not rated by this catalog.
Source and publication type
arXiv · Research preprint.
Peer-review status is not provided by this source.
Author and publication status
Author metadata is not stored; see the primary paper.
Threat model and attacker access
Black-box model, service, or application access.
Related deployment categories
No related deployment category is classified.
Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
Affected systems
Large Language Models (LLMs) such as GPT-4, Claude 2, and Vicuna, and potentially other LLMs equipped with similar safety mechanisms are affected. The vulnerability is independent of the specific model architecture or…

Research Paper

Scalable and transferable black-box jailbreaks for language models via persona modulation

Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.

View Paper

Evidence

This entry is based on a primary research source. Its findings are paper-reported; independent reproduction and verification are not claimed.

https://arxiv.org/abs/2311.03348