Skip to main content
LLM Security Database
Skip to research details
Back to research findings
LMVD-ID: 525dd110
Paper published June 1, 2024
Entry analyzed December 29, 2024
Paper-reported evidence
Confidence: Source-linked

The LMVD-ID is an internal research identifier, not an official CVE identifier.

RL-Powered LLM Jailbreak

RL-JACK is a reinforcement learning-based black-box attack that generates jailbreaking prompts to bypass safety mechanisms in LLMs. The attack leverages a deep reinforcement learning agent to iteratively refine…

BibTeX citation

Paper-evaluated models(6)

Falcon-40B-instruct, GPT-3.5 Turbo, Llama 2 70B Chat +3 more
  • Falcon-40B-instruct
  • GPT-3.5 Turbo
  • Llama 2 70B Chat
  • Llama 2 7B Chat
  • Vicuna 13B
  • Vicuna 7B

Description

RL-JACK is a reinforcement learning-based black-box attack that generates jailbreaking prompts to bypass safety mechanisms in LLMs. The attack leverages a deep reinforcement learning agent to iteratively refine prompts, maximizing the likelihood of eliciting harmful responses to unethical questions. The effectiveness stems from a novel reward function that provides continuous feedback based on cosine similarity to a reference answer from an unaligned LLM, and an action space that strategically modifies prompts using diverse techniques (e.g., creating role-playing scenarios).

Examples

See the RL-JACK paper for specific examples of generated jailbreaking prompts and their effectiveness against various LLMs. The paper includes detailed examples against models like Llama2-70b and GPT-3.5.

Impact

Successful exploitation of this vulnerability allows attackers to circumvent LLMs' safety features, potentially causing the generation of harmful, unethical, or illegal content. This can lead to the spread of misinformation, the creation of malicious software, or other damaging consequences.

Affected Systems

A wide range of LLMs are affected, including both open-source models (e.g., Llama2, Vicuna, Falcon) and commercial models (e.g., GPT-3.5). The vulnerability is demonstrated against multiple LLMs with varying levels of safety alignment.

Mitigation Steps

  • Implement robust prompt filtering and input sanitization mechanisms designed to detect and block potentially malicious prompts.
  • Develop and deploy more sophisticated safety alignment techniques that are resilient to iterative prompt manipulation.
  • Incorporate adversarial training methods to enhance the model's resilience to jailbreaking attacks during the training phase.
  • Continuously monitor and update safety mechanisms based on emerging attack techniques.
  • Investigate and consider alternative reward methods for safety training that are more robust to prompt manipulation.

Research context and confidence

Evidence and verification
Paper-reported; independent reproduction is not documented.
Primary research source linked.
Severity
Not rated by this catalog.
Source and publication type
arXiv · Research preprint.
Peer-review status is not provided by this source.
Author and publication status
Author metadata is not stored; see the primary paper.
Threat model and attacker access
Black-box model, service, or application access.
Related deployment categories
No related deployment category is classified.
Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
Affected systems
A wide range of LLMs are affected, including both open-source models (e.g., Llama2, Vicuna, Falcon) and commercial models (e.g., GPT-3.5). The vulnerability is demonstrated against multiple LLMs with varying levels of…

Research Paper

RL-JACK: Reinforcement Learning-powered Black-box Jailbreaking Attack against LLMs

Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.

View Paper

Evidence

This entry is based on a primary research source. Its findings are paper-reported; independent reproduction and verification are not claimed.

https://arxiv.org/abs/2406.08725