RoTBench: A Multi-Level Benchmark for Evaluating the Robustness of Large Language Models in Tool Learning
Junjie Ye, Yilong Wu, Songyang Gao, Caishuang Huang, Sixian Li, Guanyu Li, Xiaoran Fan, Qi Zhang, Tao Gui, Xuanjing Huang
摘要
Tool learning has generated widespread interest as a vital means of interaction between Large Language Models (LLMs) and the physical world. Current research predominantly emphasizes LLMs' capacity to utilize tools in well-structured environments while overlooking their stability when confronted with the inevitable noise of the real world. To bridge this gap, we introduce RoTBench, a multi-level benchmark for evaluating the robustness of LLMs in tool learning. Specifically, we establish five external environments, each featuring varying levels of noise (i.e., Clean, Slight, Medium, Heavy, and Union), providing an in-depth analysis of the model's resilience across three critical phases: tool selection, parameter identification, and content filling. Experiments involving six widely-used models underscore the urgent necessity for enhancing the robustness of LLMs in tool learning. For instance, the performance of GPT-4 even drops significantly from 80.00 to 58.10 when there is no substantial change in manual accuracy. More surprisingly, the noise correction capability inherent in the GPT family paradoxically impedes its adaptability in the face of mild noise. In light of these findings, we propose RoTTuning, a strategy that enriches the diversity of training environments to bolster the robustness of LLMs in tool learning. The code and data are available at https: //github.com/Junjie-Ye/RoTBench .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing AgentsMuyu He, Anand Kumar, Soumyadeep Bakshi, James Zou 等ACL 2026 · 被引用 8 次
- VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language ModelsRohit Saxena, Alessandro Suglia, Pasquale MinerviniICML 2026 · 被引用 7 次
- Trustworthy Medical Question Answering: An Evaluation-Centric SurveyYinuo Wang, Baiyang Wang, Robert E. Mercer, Frank Rudzicz 等EMNLP 2025 · 被引用 2 次
- reWordBench: Benchmarking and Improving the Robustness of Reward Models with Transformed InputsZhaofeng Wu, Michihiro Yasunaga, Andrew Cohen, Yoon Kim 等EMNLP 2025
- LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language ModelsMing Zhang, Yujiong Shen, Jingyi Deng, Yuhui Wang 等ACL 2026
它引用的顶会 Paper7
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
- ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool EmbeddingsShibo Hao, Tianyang Liu, Zhen Wang, Zhiting HuNeurIPS 2023 · 被引用 315 次
- MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to UseYue Huang, Jiawen Shi, Yuan Li, Chenrui Fan 等ICLR 2024 · 被引用 188 次
- Towards Robust and Safe Reinforcement Learning with Benign Off-policy DataZuxin Liu, Zijian Guo, Zhepeng Cen, Huan Zhang 等ICML 2023 · 被引用 14 次
相关 Paper
- MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language ModelsPei Wang, Yanan Wu, Noah Wang, Jiaheng Liu 等ICLR 2025
- AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy ConditionRuipeng Wang, Yuxin Chen, Yukai Wang, Chang Wu 等ICML 2026 · 被引用 12 次
- ResiliBench: Evaluating Agentic Workflow Adaptation in Stochastic EnvironmentsRuicheng Ao, Zeping Min, Tingyu Zhu, Wotao Yin 等ICLR 2026
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu 等ICLR 2024 · 被引用 1,469 次
- MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language FeedbackXingyao Wang, Zihan Wang, Jiateng Liu, Yangyi Chen 等ICLR 2024 · 被引用 308 次
