Subtle Risks, Critical Failures: A Framework for Diagnosing Physical Safety of LLMs for Embodied Decision Making
Yejin Son, Minseo Kim, Sungwoong Kim, Seungju Han, Jian Kim, Dongju Jang, Youngjae Yu, Chan Young Park
Abstract
Large Language Models (LLMs) are increasingly used for decision making in embodied agents, yet existing safety evaluations often rely on coarse success rates and domain-specific setups, making it difficult to diagnose why and where these models fail. This obscures our understanding of embodied safety and limits the selective deployment of LLMs in high-risk physical environments. We introduce SAFEL, the framework for systematically evaluating the physical safety of LLMs in embodied decision making. SAFEL assesses two key competencies: (1) rejecting unsafe commands via the Command Refusal Test, and (2) generating safe and executable plans via the Plan Safety Test. Critically, the latter is decomposed into functional modules, goal interpretation, transition modeling, action sequencing, enabling fine-grained diagnosis of safety failures. To support this framework, we introduce EMBODYGUARD, a PDDL-grounded benchmark containing 942 LLM-generated scenarios covering both overtly malicious and contextually hazardous instructions. Evaluation across 13 state-of-the-art LLMs reveals that while models often reject clearly unsafe commands, they struggle to anticipate and mitigate subtle, situational risks. Our results highlight critical limitations in current LLMs and provide a foundation for more targeted, modular improvements in safe embodied reasoning. * (holding ?knife) next to a human * (inside ?human freezer.n.01_1) * (toggled_on stove.n.01_1) with flammable objects nearby * (nextto ?agent ?human) with a harmful object -Output Format: Refusal, o1, [brief explanation]
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 89621d36-5c15-4b4a-abf7-ac261e19705eCited by top-tier papers3
- IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household TasksXiaoya Lu, Zeren Chen, Xuhao Hu, Yijin Zhou et al.AAAI 2026 · 20 citations
- Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark StudyDongGeon Lee, Joonwon Jang, Jihae Jeong, Hwanjo YuEMNLP 2025 · 1 citation
- EMBGuard: Constructing Hazard-Aware Guardrails for Safe Planning in Embodied AgentsDongwook Choi, Taeyoon Kwon, Bogyung Jeong, Minju Kim et al.ICML 2026
Builds on3
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao et al.ICCV 2023 · 685 citations
- Identifying the Risks of LM Agents with an LM-Emulated SandboxYangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis et al.ICLR 2024 · 292 citations
- Safe Planner: Empowering Safety Awareness in Large Pre-Trained Models for Robot Task PlanningSiyuan Li, Feifan Liu, Lingfei Cui, Jiani Lu et al.AAAI 2025 · 4 citations
Related papers
- AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous InstructionsZonghao Ying, Le Wang, Yisong Xiao, Jiakai Wang et al.CVPR 2026 · 42 citations
- Multimodal Situational SafetyKaiwen Zhou, Chengzhi Liu, Xuandong Zhao, Anderson Compalas et al.ICLR 2025
- HAZARD Challenge: Embodied Decision Making in Dynamically Changing EnvironmentsQinhong Zhou, Sunli Chen, Yisong Wang, Haozhe Xu et al.ICLR 2024 · 33 citations
- SafeText: A Benchmark for Exploring Physical Safety in Language ModelsSharon Levy, Emily Allaway, Melanie Subbiah, Lydia B. Chilton et al.EMNLP 2022 · 14 citations
- EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied AgentsRui Yang, Hanyang Chen, Junyu Zhang, Mark Zhao et al.ICML 2025
