Identifying the Risks of LM Agents with an LM-Emulated Sandbox
Yangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis, Yongchao Zhou, Jimmy Ba, Yann Dubois, Chris J. Maddison, Tatsunori Hashimoto
Abstract
Recent advances in Language Model (LM) agents and tool use, exemplified by applications like ChatGPT Plugins, enable a rich set of capabilities but also amplify potential risks - such as leaking private data or causing financial losses. Identifying these risks is labor-intensive, necessitating implementing the tools, setting up the environment for each test scenario manually, and finding risky cases. As tools and agents become more complex, the high cost of testing these agents will make it increasingly difficult to find high-stakes, long-tailed risks. To address these challenges, we introduce ToolEmu: a framework that uses an LM to emulate tool execution and enables the testing of LM agents against a diverse range of tools and scenarios, without manual instantiation. Alongside the emulator, we develop an LM-based automatic safety evaluator that examines agent failures and quantifies associated risks. We test both the tool emulator and evaluator through human evaluation and find that 68.8% of failures identified with ToolEmu would be valid real-world agent failures. Using our curated initial benchmark consisting of 36 high-stakes tools and 144 test cases, we provide a quantitative risk analysis of current LM agents and identify numerous failures with potentially severe outcomes. Notably, even the safest LM agent exhibits such failures 23.9% of the time according to our evaluator, underscoring the need to develop safer LM agents for real-world deployment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0b5c974c-470e-4de3-9440-472357dec7edCited by top-tier papers72
- -Bench: Evaluating Conversational Agents in a Dual-Control EnvironmentVictor Barres, Honghua Dong, Soham Ray, Xujie Si et al.ICML 2026 · 399 citations
- OpenCUA: Open Foundations for Computer-Use AgentsXinyuan Wang, Bowen Wang, Dunjie Lu, Junlin Yang et al.NeurIPS 2025 · 151 citations
- An LLM Compiler for Parallel Function CallingSehoon Kim, Suhong Moon, Ryan Tabrizi, Nicholas Lee et al.ICML 2024 · 142 citations
- Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially FastXiangming Gu, Xiaosen Zheng, Tianyu Pang, Chao Du et al.ICML 2024 · 128 citations
- Attacking Vision-Language Computer Agents via Pop-upsYanzhe Zhang, Tao Yu, Diyi YangACL 2025 · 99 citations
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
Related papers
- AgentSentinel: An End-to-End and Real-Time Security Defense Framework for Computer-Use AgentsHaitao Hu, Peng Chen, Yanpeng Zhao, Yuqi ChenCCS 2025
- Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using AgentsXu Li, Simon Yu, Minzhou Pan, Yiyou Sun et al.ICML 2026 · 16 citations
- OpenAgentSafety: A Comprehensive Framework For Evaluating Real-World AI Agent SafetySanidhya Vijayvargiya, Aditya Bharat Soni, Xuhui Zhou, Zora Zhiruo Wang et al.ICLR 2026 · 75 citations
- Les Dissonances: Cross-Tool Harvesting and Polluting in Pool-of-Tools Empowered LLM AgentsZichuan Li, Jian Cui, Xiaojing Liao, Luyi XingNDSS 2026 · 24 citations
- Dynamic Evaluation with Cognitive Reasoning for Multi-turn Safety of Large Language ModelsLanxue Zhang, Yanan Cao, Yuqiang Xie, Fang Fang et al.ACL 2025
