Benchmarking LLM Tool-Use in the Wild
Peijie Yu, Wei Liu, Yifan Yang, Jinjian Li, Zelong Zhang, Xiao Feng, Feng Zhang
摘要
Fulfilling user needs through Large Language Model multi-turn, multi-step tool-use is rarely a straightforward process. Real user interactions are inherently , being intricate, messy, and flexible. We identify three key challenges from user behaviour: that demand efficient orchestration of tool-call topologies, spread across dialogue turns that require contextual inference, and , which mixes task queries, clarifications, and casual conversation, forcing LLMs to adjust their policies on the fly. Existing benchmarks overlook these behaviors, making the apparent progress of LLMs on tool-use spurious. To address this, we introduce , an LLM tool-use benchmark grounded in real-world user behavior patterns. Comprehensive evaluations of 57 LLMs reveal that no model achieves an accuracy of more than 15%, indicating a substantial gap in the robustness of LLMs' agentic ability. Controlled experiments and in-depth analyses further indicate that the real challenge for LLM tool-use lies not in artificially complex tasks, but in the wild nature of user behavior, emphasizing the need to reconsider the interactions among , , and .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to UseYue Huang, Jiawen Shi, Yuan Li, Chenrui Fan 等ICLR 2024 · 被引用 188 次
- AnyTool: Self-Reflective, Hierarchical Agents for Large-Scale API CallsYu Du, Fangyun Wei, Hongyang ZhangICML 2024 · 被引用 104 次
- The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMsAkshit Sinha, Arvindh Arun, Shashwat Goel, Steffen Staab 等ICLR 2026 · 被引用 64 次
- T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by StepZehui Chen, Weihua Du, Wenwei Zhang, Kuikun Liu 等ACL 2024 · 被引用 7 次
- Direct Multi-Turn Preference Optimization for Language AgentsWentao Shi, Mengqi Yuan, Junkang Wu, Qifan Wang 等EMNLP 2024 · 被引用 1 次
相关 Paper
- BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex InstructionsTerry Yue Zhuo, Minh Chien Vu, Jenny Chim, Han Hu 等ICLR 2025
- MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language ModelsPei Wang, Yanan Wu, Noah Wang, Jiaheng Liu 等ICLR 2025
- OrchestrationBench: LLM-Driven Agentic Planning and Tool Use in Multi-Domain ScenariosAelim Ahn, Sooyeon Lee, Hyosun Wang, Chiwan Park 等ICLR 2026
- CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World UncertaintyJohannes Kirmayr, Lukas Stappen, Elisabeth AndréACL 2026 · 被引用 5 次
- Learning to Ask: When LLM Agents Meet Unclear InstructionWenxuan Wang, Juluan Shi, Zixuan Ling, Yuk-Kit Chan 等EMNLP 2025 · 被引用 1 次
