The LMVD-ID is an internal research identifier, not an official CVE identifier.
LLM Robot Bias & Violence
Large Language Models (LLMs) used to control robots exhibit biases leading to discriminatory and unsafe behaviors. When provided with personal characteristics (e.g., race, gender, disability), LLMs generate biased…
Paper-evaluated models(4)
- GPT-3.5
- GPT-3.5 Turbo
- GPT-4
- Mistral 7B
Research context and confidence
- Evidence and verification
- Paper-reported; independent reproduction is not documented.
- Primary research source linked.
- Severity
- Not rated by this catalog.
- Source and publication type
- arXiv · Research preprint.
- Peer-review status is not provided by this source.
- Author and publication status
- Author metadata is not stored; see the primary paper.
- Threat model and attacker access
- Black-box model, service, or application access.
- Related deployment categories
- No related deployment category is classified.
- Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
- Affected systems
- Robotic systems utilizing LLMs for decision-making, task planning, and human interaction, regardless of vendor. Specific LLMs affected include, but are not limited to, GPT-3.5, Mistral 7b v0.1, Gemini, CoPilot (powered…
Research Paper
Llm-driven robots risk enacting discrimination, violence, and unlawful actions
Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.
View PaperDescription
Large Language Models (LLMs) used to control robots exhibit biases leading to discriminatory and unsafe behaviors. When provided with personal characteristics (e.g., race, gender, disability), LLMs generate biased outputs resulting in discriminatory actions (e.g., assigning lower rescue priority to certain groups) and accept or deem feasible dangerous or unlawful instructions (e.g., removing a person's mobility aid).
Examples
See the paper for numerous examples demonstrating discriminatory outputs and acceptance of harmful instructions across various LLM models and prompting scenarios. Specific examples include assigning low rescue priority to certain ethnicities and accepting instructions to remove a person's mobility aid.
Impact
LLM-driven robots may enact discrimination, violence, and unlawful actions, resulting in physical and psychological harm to individuals, particularly those belonging to marginalized groups. This poses significant safety and ethical concerns.
Affected Systems
Robotic systems utilizing LLMs for decision-making, task planning, and human interaction, regardless of vendor. Specific LLMs affected include, but are not limited to, GPT-3.5, Mistral 7b v0.1, Gemini, CoPilot (powered by GPT-4), and Llama 2.
Mitigation Steps
- Implement rigorous bias detection and mitigation techniques during LLM training and deployment.
- Develop robust safety frameworks and filters to prevent the execution of harmful instructions.
- Conduct comprehensive risk assessments of LLM-driven robotic systems, incorporating evaluations across various demographic groups and scenarios.
- Prioritize human oversight and intervention mechanisms to ensure safe and responsible robot operation. Limit the capabilities of LLMs in high-risk scenarios.
- Focus on validation within specific Operational Design Domains (ODDs) rather than aiming for general-purpose safety.
Evidence
This entry is based on a primary research source. Its findings are paper-reported; independent reproduction and verification are not claimed.
https://arxiv.org/abs/2406.08824Related research
- Voice Agent Behavioral Bypass
Published February 1, 2026 · model-layer, application-layer, injection
- LLM Hate Campaign Vulnerability
Published January 1, 2025 · application-layer, injection, extraction
- AI Browser Indirect Injection
Published October 1, 2025 · application-layer, prompt-layer, injection