Subtle Risks, Critical Failures: A Framework for Diagnosing Physical Safety of LLMs for Embodied Decision Making
Yejin Son, Minseo Kim, Sungwoong Kim, Seungju Han, Jian Kim, Dongju Jang, Youngjae Yu, Chan Young Park
摘要
Large Language Models (LLMs) are increasingly used for decision making in embodied agents, yet existing safety evaluations often rely on coarse success rates and domain-specific setups, making it difficult to diagnose why and where these models fail. This obscures our understanding of embodied safety and limits the selective deployment of LLMs in high-risk physical environments. We introduce SAFEL, the framework for systematically evaluating the physical safety of LLMs in embodied decision making. SAFEL assesses two key competencies: (1) rejecting unsafe commands via the Command Refusal Test, and (2) generating safe and executable plans via the Plan Safety Test. Critically, the latter is decomposed into functional modules, goal interpretation, transition modeling, action sequencing, enabling fine-grained diagnosis of safety failures. To support this framework, we introduce EMBODYGUARD, a PDDL-grounded benchmark containing 942 LLM-generated scenarios covering both overtly malicious and contextually hazardous instructions. Evaluation across 13 state-of-the-art LLMs reveals that while models often reject clearly unsafe commands, they struggle to anticipate and mitigate subtle, situational risks. Our results highlight critical limitations in current LLMs and provide a foundation for more targeted, modular improvements in safe embodied reasoning. * (holding ?knife) next to a human * (inside ?human freezer.n.01_1) * (toggled_on stove.n.01_1) with flammable objects nearby * (nextto ?agent ?human) with a harmful object -Output Format: Refusal, o1, [brief explanation]
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household TasksXiaoya Lu, Zeren Chen, Xuhao Hu, Yijin Zhou 等AAAI 2026 · 被引用 20 次
- Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark StudyDongGeon Lee, Joonwon Jang, Jihae Jeong, Hwanjo YuEMNLP 2025 · 被引用 1 次
- EMBGuard: Constructing Hazard-Aware Guardrails for Safe Planning in Embodied AgentsDongwook Choi, Taeyoon Kwon, Bogyung Jeong, Minju Kim 等ICML 2026
它引用的顶会 Paper3
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao 等ICCV 2023 · 被引用 685 次
- Identifying the Risks of LM Agents with an LM-Emulated SandboxYangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis 等ICLR 2024 · 被引用 292 次
- Safe Planner: Empowering Safety Awareness in Large Pre-Trained Models for Robot Task PlanningSiyuan Li, Feifan Liu, Lingfei Cui, Jiani Lu 等AAAI 2025 · 被引用 4 次
相关 Paper
- AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous InstructionsZonghao Ying, Le Wang, Yisong Xiao, Jiakai Wang 等CVPR 2026 · 被引用 42 次
- Multimodal Situational SafetyKaiwen Zhou, Chengzhi Liu, Xuandong Zhao, Anderson Compalas 等ICLR 2025
- HAZARD Challenge: Embodied Decision Making in Dynamically Changing EnvironmentsQinhong Zhou, Sunli Chen, Yisong Wang, Haozhe Xu 等ICLR 2024 · 被引用 33 次
- SafeText: A Benchmark for Exploring Physical Safety in Language ModelsSharon Levy, Emily Allaway, Melanie Subbiah, Lydia B. Chilton 等EMNLP 2022 · 被引用 14 次
- EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied AgentsRui Yang, Hanyang Chen, Junyu Zhang, Mark Zhao 等ICML 2025
