RoTBench: A Multi-Level Benchmark for Evaluating the Robustness of Large Language Models in Tool Learning
Junjie Ye, Yilong Wu, Songyang Gao, Caishuang Huang, Sixian Li, Guanyu Li, Xiaoran Fan, Qi Zhang, Tao Gui, Xuanjing Huang
Abstract
Tool learning has generated widespread interest as a vital means of interaction between Large Language Models (LLMs) and the physical world. Current research predominantly emphasizes LLMs' capacity to utilize tools in well-structured environments while overlooking their stability when confronted with the inevitable noise of the real world. To bridge this gap, we introduce RoTBench, a multi-level benchmark for evaluating the robustness of LLMs in tool learning. Specifically, we establish five external environments, each featuring varying levels of noise (i.e., Clean, Slight, Medium, Heavy, and Union), providing an in-depth analysis of the model's resilience across three critical phases: tool selection, parameter identification, and content filling. Experiments involving six widely-used models underscore the urgent necessity for enhancing the robustness of LLMs in tool learning. For instance, the performance of GPT-4 even drops significantly from 80.00 to 58.10 when there is no substantial change in manual accuracy. More surprisingly, the noise correction capability inherent in the GPT family paradoxically impedes its adaptability in the face of mild noise. In light of these findings, we propose RoTTuning, a strategy that enriches the diversity of training environments to bolster the robustness of LLMs in tool learning. The code and data are available at https: //github.com/Junjie-Ye/RoTBench .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing AgentsMuyu He, Anand Kumar, Soumyadeep Bakshi, James Zou et al.ACL 2026 · 8 citations
- VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language ModelsRohit Saxena, Alessandro Suglia, Pasquale MinerviniICML 2026 · 7 citations
- Trustworthy Medical Question Answering: An Evaluation-Centric SurveyYinuo Wang, Baiyang Wang, Robert E. Mercer, Frank Rudzicz et al.EMNLP 2025 · 2 citations
- reWordBench: Benchmarking and Improving the Robustness of Reward Models with Transformed InputsZhaofeng Wu, Michihiro Yasunaga, Andrew Cohen, Yoon Kim et al.EMNLP 2025
- LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language ModelsMing Zhang, Yujiong Shen, Jingyi Deng, Yuhui Wang et al.ACL 2026
Builds on7
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool EmbeddingsShibo Hao, Tianyang Liu, Zhen Wang, Zhiting HuNeurIPS 2023 · 315 citations
- MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to UseYue Huang, Jiawen Shi, Yuan Li, Chenrui Fan et al.ICLR 2024 · 188 citations
- Towards Robust and Safe Reinforcement Learning with Benign Off-policy DataZuxin Liu, Zijian Guo, Zhepeng Cen, Huan Zhang et al.ICML 2023 · 14 citations
Related papers
- MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language ModelsPei Wang, Yanan Wu, Noah Wang, Jiaheng Liu et al.ICLR 2025
- AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy ConditionRuipeng Wang, Yuxin Chen, Yukai Wang, Chang Wu et al.ICML 2026 · 12 citations
- ResiliBench: Evaluating Agentic Workflow Adaptation in Stochastic EnvironmentsRuicheng Ao, Zeping Min, Tingyu Zhu, Wotao Yin et al.ICLR 2026
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu et al.ICLR 2024 · 1,469 citations
- MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language FeedbackXingyao Wang, Zihan Wang, Jiateng Liu, Yangyi Chen et al.ICLR 2024 · 308 citations
