Skip to main content
LLM Security Database
Skip to research details
Back to research findings
LMVD-ID: 6b3d9d19
Paper published March 1, 2025
Entry analyzed January 14, 2026
Paper-reported evidence
Confidence: Source-linked

The LMVD-ID is an internal research identifier, not an official CVE identifier.

Cyber Fine-Tuning Undermines Safety

Fine-tuning Large Language Models (LLMs) on the CyberLLMInstruct dataset results in a critical degradation of safety alignment and refusal mechanisms. While the dataset comprises "pseudo-malicious" content (educational…

BibTeX citation

Paper-evaluated models(7)

Llama 2 70B, Llama 3 8B, Llama 3.1 8B +4 more
  • Llama 2 70B
  • Llama 3 8B
  • Llama 3.1 8B
  • Phi-3 Mini 3.8B
  • Mistral 7B
  • Qwen 2.5 7B
  • Gemma 2 9B

Description

Fine-tuning Large Language Models (LLMs) on the CyberLLMInstruct dataset results in a critical degradation of safety alignment and refusal mechanisms. While the dataset comprises "pseudo-malicious" content (educational descriptions of malware, phishing, and exploits without executable payloads), the Supervised Fine-Tuning (SFT) process on this corpus causes the models to generalize this instruction-following behavior to actual malicious requests. This effectively bypasses safety guardrails established during pre-training. Specifically, fine-tuned models exhibit a drastic reduction in resilience against Prompt Injection and Sensitive Information Disclosure (OWASP Top 10 for LLMs), allowing attackers to generate actionable malware descriptions, social engineering templates, and vulnerability exploits that the base models would otherwise refuse.

Examples

The vulnerability is reproducible by fine-tuning standard open-source models on the provided dataset and subjecting them to adversarial testing.

  1. Fine-tune Llama 3.1 8B on the CyberLLMInstruct dataset using standard SFT (10 epochs, learning rate 2e-4).
  2. Attempt a standard prompt injection attack requesting a phishing email template or malware logic.
  3. Result: The base model achieves a security score of 0.95 (refusal). The fine-tuned model's score drops to 0.15, generating the malicious content.
  • Attack Vectors: The dataset trains the model to comply with requests regarding:
  • Malware creation (ransomware, Trojans, spyware logic).
  • Social engineering (phishing campaign structures).
  • Zero-day exploit descriptions.

Impact

  • Safety Bypass: Complete or near-complete removal of safety guardrails regarding cyber-offensive content.
  • Malware Generation: Automated generation of sophisticated phishing campaigns, malicious scripts, and exploit methodologies.
  • Democratization of Attacks: Lowers the entry barrier for malicious actors to utilize high-performing LLMs for offensive cyber operations (Crime-as-a-Service).
  • Vulnerability Exposure: Increases model susceptibility to Prompt Injection, Sensitive Information Disclosure, and Data Poisoning.

Affected Systems

  • Dataset: CyberLLMInstruct (v1.0 and associated releases).
  • Models: Any LLM fine-tuned on this dataset, confirmed specifically on:
  • Llama 3 / 3.1 (8B)
  • Llama 2 (70B)
  • Phi 3 Mini (3.8B)
  • Mistral 7B
  • Qwen 2.5 7B
  • Gemma 2 9B

Mitigation Steps

  • Restricted Access: Limit the deployment of models fine-tuned on CyberLLMInstruct to controlled environments accessible only by verified domain experts (red teams/pen-testers).
  • Safety-Aware Fine-Tuning: Do not use the dataset for SFT without incorporating a concurrent safety alignment dataset (e.g., refusal samples) to preserve guardrails (Mixed Fine-Tuning).
  • Post-Training Red Teaming: Subject fine-tuned models to rigorous adversarial testing frameworks (e.g., DeepEval, OWASP Top 10 benchmarks) before any release.
  • Human-in-the-Loop: Ensure all outputs from models trained on this data are reviewed by security professionals to filter erroneous or dangerous suggestions.

Research context and confidence

Evidence and verification
Paper-reported; independent reproduction is not documented.
Primary research source linked.
Severity
Not rated by this catalog.
Source and publication type
arXiv · Research preprint.
Peer-review status is not provided by this source.
Author and publication status
Author metadata is not stored; see the primary paper.
Threat model and attacker access
White-box access to model or deployment internals.
Related deployment categories
Fine-tuning
Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
Affected systems
Dataset: CyberLLMInstruct (v1.0 and associated releases). Models: Any LLM fine-tuned on this dataset, confirmed specifically on: Llama 3 / 3.1 (8B) Llama 2 (70B) Phi 3 Mini (3.8B) Mistral 7B Qwen 2.5 7B Gemma 2 9B

Research Paper

CyberLLMInstruct: A new dataset for analysing safety of fine-tuned LLMs using cyber security data

Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.

View Paper

Evidence

This entry is based on a primary research source. Its findings are paper-reported; independent reproduction and verification are not claimed.

https://arxiv.org/abs/2503.09334