AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions
Zonghao Ying, Le Wang, Yisong Xiao, Jiakai Wang, Yuqing Ma, Jinyang Guo, Zhenfei Yin, Mingchuan Zhang, Aishan Liu, Xianglong Liu
Abstract
The integration of vision-language models (VLMs) is driving a new generation of embodied agents capable of operating in human-centered environments. However, as deployment expands, these systems face growing safety risks, particularly when executing hazardous instructions. Current safety evaluation benchmarks remain limited: they cover only narrow scopes of hazards and focus primarily on final outcomes, neglecting the agent's full perception-planning-execution process and thereby obscuring critical failure modes. Therefore, we present SAFE, a benchmark for systematically assessing the safety of embodied VLM agents on hazardous instructions. SAFE comprises three components: SAFE-THOR, an extensible adversarial simulation sandbox with a universal adapter that maps high-level VLM outputs to low-level embodied controls, supporting diverse agent workflow integration; SAFE-VERSE, a risk-aware task suite inspired by Asimov's Three Laws of Robotics, comprising 45 adversarial scenarios, 1,350 hazardous tasks, and 9,900 instructions that span risks to humans, environments, and agents; and SAFE-DIAGNOSE, a multi-level and fine-grained evaluation protocol measuring agent performance across perception, planning, and execution. Applying SAFE to nine state-of-the-art VLMs and two embodied agent workflows, we uncover systematic failures in translating hazard recognition into safe planning and execution. Our findings reveal fundamental limitations in current safety alignment and demonstrate the necessity of a comprehensive, multi-stage evaluation for developing safer embodied intelligence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 59901f11-149a-4471-9690-206d613d41d2Cited by top-tier papers1
Ask how each one uses itBuilds on10
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 2,230 citations
- ExpeL: LLM Agents Are Experiential LearnersAndrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin et al.AAAI 2024 · 484 citations
- GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via CipherYouliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang et al.ICLR 2024 · 441 citations
- Multilingual Jailbreak Challenges in Large Language ModelsYue Deng, Wenxuan Zhang, Sinno Jialin Pan, Lidong BingICLR 2024 · 230 citations
- How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMsYi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang et al.ACL 2024 · 64 citations
Related papers
- Subtle Risks, Critical Failures: A Framework for Diagnosing Physical Safety of LLMs for Embodied Decision MakingYejin Son, Minseo Kim, Sungwoong Kim, Seungju Han et al.EMNLP 2025 · 7 citations
- IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household TasksXiaoya Lu, Zeren Chen, Xuhao Hu, Yijin Zhou et al.AAAI 2026 · 20 citations
- How Foundational Skills Influence VLM-based Embodied Agents: A Native PerspectiveBo Peng, Pi Bu, Keyu Pan, Xinrun Xu et al.AAAI 2026 · 1 citation
- OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic WorkflowsQiushi Sun, Mukai Li, Zhoumianze Liu, Zhihui Xie et al.ACL 2026 · 14 citations
- EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied AgentsRui Yang, Hanyang Chen, Junyu Zhang, Mark Zhao et al.ICML 2025
