Published 10/1/2023
Analyzed 12/28/2024
Large Language Models (LLMs) employing Reinforcement Learning from Human Feedback (RLHF) and instruction tuning methods may exhibit superficial safety guardrails vulnerable to parametric red-teaming attacks. Fine-tuning the model on a dataset of harmful prompts and their corresponding helpful (but harmful) responses can bypass built-in safety mechanisms, resulting in the model generating unsafe outputs. This vulnerability is demonstrated by achieving an 88% success rate in eliciting harmful…
Language model unalignment: Parametric red-teaming to expose hidden harms and biases
Evaluated models: Claude 1, Claude 2, GPT-4 +6 more
Source: arXiv