Skip to main content
LLM Security Database
Skip to research details
Back to Vulnerability List
LMVD-ID: e7254d62
Paper published June 1, 2024
Entry analyzed April 12, 2025
Severity: Not rated
Confidence: Source-linked

The LMVD-ID is an internal research identifier, not an official CVE identifier.

LLM Robot Bias & Violence

Large Language Models (LLMs) used to control robots exhibit biases leading to discriminatory and unsafe behaviors. When provided with personal characteristics (e.g., race, gender, disability), LLMs generate biased…

BibTeX citation

Paper-evaluated models(4)

  • GPT-3.5
  • GPT-3.5 Turbo
  • GPT-4
  • Mistral 7B

Research context and confidence

Evidence and verification
Paper-reported; independent reproduction is not documented.
Primary research source linked.
Severity
Not rated by this catalog.
Source and publication type
arXiv · Research preprint.
Peer-review status is not provided by this source.
Author and publication status
Author metadata is not stored; see the primary paper.
Threat model and attacker access
Black-box model, service, or application access.
Related deployment categories
No related deployment category is classified.
Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
Affected systems
Robotic systems utilizing LLMs for decision-making, task planning, and human interaction, regardless of vendor. Specific LLMs affected include, but are not limited to, GPT-3.5, Mistral 7b v0.1, Gemini, CoPilot (powered…

Research Paper

Llm-driven robots risk enacting discrimination, violence, and unlawful actions

Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.

View Paper

Description

Large Language Models (LLMs) used to control robots exhibit biases leading to discriminatory and unsafe behaviors. When provided with personal characteristics (e.g., race, gender, disability), LLMs generate biased outputs resulting in discriminatory actions (e.g., assigning lower rescue priority to certain groups) and accept or deem feasible dangerous or unlawful instructions (e.g., removing a person's mobility aid).

Examples

See the paper for numerous examples demonstrating discriminatory outputs and acceptance of harmful instructions across various LLM models and prompting scenarios. Specific examples include assigning low rescue priority to certain ethnicities and accepting instructions to remove a person's mobility aid.

Impact

LLM-driven robots may enact discrimination, violence, and unlawful actions, resulting in physical and psychological harm to individuals, particularly those belonging to marginalized groups. This poses significant safety and ethical concerns.

Affected Systems

Robotic systems utilizing LLMs for decision-making, task planning, and human interaction, regardless of vendor. Specific LLMs affected include, but are not limited to, GPT-3.5, Mistral 7b v0.1, Gemini, CoPilot (powered by GPT-4), and Llama 2.

Mitigation Steps

  • Implement rigorous bias detection and mitigation techniques during LLM training and deployment.
  • Develop robust safety frameworks and filters to prevent the execution of harmful instructions.
  • Conduct comprehensive risk assessments of LLM-driven robotic systems, incorporating evaluations across various demographic groups and scenarios.
  • Prioritize human oversight and intervention mechanisms to ensure safe and responsible robot operation. Limit the capabilities of LLMs in high-risk scenarios.
  • Focus on validation within specific Operational Design Domains (ODDs) rather than aiming for general-purpose safety.

Evidence

This entry is based on a primary research source. Its findings are paper-reported; independent reproduction and verification are not claimed.

https://arxiv.org/abs/2406.08824