Tools Fail: Detecting Silent Errors in Faulty Tools
Jimin Sun, So Yeon Min, Yingshan Chang, Yonatan Bisk
Abstract
Tools have become a mainstay of LLMs, allowing them to retrieve knowledge not in their weights, to perform tasks on the web, and even to control robots. However, most ontologies and surveys of tool-use have assumed the core challenge for LLMs is choosing the tool. Instead, we introduce a framework for tools more broadly which guides us to explore a model’s ability to detect “silent” tool errors, and reflect on how to plan. This more directly aligns with the increasingly popular use of models as tools. We provide an initial approach to failure recovery with promising results both on a controlled calculator setting and embodied agent planning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web AgentsIdo Levy, Ben wiesel, Sami Marreed, Alon Oved et al.ICLR 2026 · 78 citations
- Towards Reliable Benchmarking: A Contamination Free, Controllable Evaluation Framework for Multi-step LLM Function CallingSeiji Maekawa, Jackson Hassell, Pouya Pezeshkpour, Tom M. Mitchell et al.ICLR 2026 · 14 citations
- ET-Agent: Incentivizing Effective Tool-Integrated Reasoning Agent via Behavior CalibrationYifei Chen, Guanting Dong, Zhicheng DouACL 2026 · 3 citations
- CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error ScenariosShiting Huang, Zhen Fang, Zehui Chen, Siyu Yuan et al.EMNLP 2025
- Grounding LLMs in Scientific Discovery via Embodied ActionsBo Zhang, Jinfeng Zhou, Yuxuan Chen, Jianing Yin et al.ICML 2026
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
Related papers
- Adaptive Tool Use in Large Language Models with Meta-Cognition TriggerWenjun Li, Dexun Li, Kuicai Dong, Cong Zhang et al.ACL 2025
- ToolOmni: Enabling Open-World Tool Use via Agentic learning with Proactive Retrieval and Grounded ExecutionShouzheng Huang, Meishan Zhang, Baotian Hu, Min ZhangACL 2026
- Predicting Emergent Tool Use in LLMs Before It Emerges: A Proxy PerspectiveBowen Zhang, Yan Yan, Guang Liu, Xu-Cheng YinAAAI 2026
- The Confidence Dichotomy: Analyzing and Mitigating Miscalibration in Tool-Use AgentsWeihao Xuan, Qingcheng Zeng, Heli Qi, Yunze Xiao et al.ACL 2026 · 4 citations
- Learning to Ask: When LLM Agents Meet Unclear InstructionWenxuan Wang, Juluan Shi, Zixuan Ling, Yuk-Kit Chan et al.EMNLP 2025 · 1 citation
