Benchmarking LLM Tool-Use in the Wild
Peijie Yu, Wei Liu, Yifan Yang, Jinjian Li, Zelong Zhang, Xiao Feng, Feng Zhang
Abstract
Fulfilling user needs through Large Language Model multi-turn, multi-step tool-use is rarely a straightforward process. Real user interactions are inherently , being intricate, messy, and flexible. We identify three key challenges from user behaviour: that demand efficient orchestration of tool-call topologies, spread across dialogue turns that require contextual inference, and , which mixes task queries, clarifications, and casual conversation, forcing LLMs to adjust their policies on the fly. Existing benchmarks overlook these behaviors, making the apparent progress of LLMs on tool-use spurious. To address this, we introduce , an LLM tool-use benchmark grounded in real-world user behavior patterns. Comprehensive evaluations of 57 LLMs reveal that no model achieves an accuracy of more than 15%, indicating a substantial gap in the robustness of LLMs' agentic ability. Controlled experiments and in-depth analyses further indicate that the real challenge for LLM tool-use lies not in artificially complex tasks, but in the wild nature of user behavior, emphasizing the need to reconsider the interactions among , , and .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 103ba8c3-2e36-458f-9ecf-41a20172ef88Builds on8
- MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to UseYue Huang, Jiawen Shi, Yuan Li, Chenrui Fan et al.ICLR 2024 · 188 citations
- AnyTool: Self-Reflective, Hierarchical Agents for Large-Scale API CallsYu Du, Fangyun Wei, Hongyang ZhangICML 2024 · 104 citations
- The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMsAkshit Sinha, Arvindh Arun, Shashwat Goel, Steffen Staab et al.ICLR 2026 · 64 citations
- T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by StepZehui Chen, Weihua Du, Wenwei Zhang, Kuikun Liu et al.ACL 2024 · 7 citations
- Direct Multi-Turn Preference Optimization for Language AgentsWentao Shi, Mengqi Yuan, Junkang Wu, Qifan Wang et al.EMNLP 2024 · 1 citation
Related papers
- BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex InstructionsTerry Yue Zhuo, Minh Chien Vu, Jenny Chim, Han Hu et al.ICLR 2025
- MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language ModelsPei Wang, Yanan Wu, Noah Wang, Jiaheng Liu et al.ICLR 2025
- OrchestrationBench: LLM-Driven Agentic Planning and Tool Use in Multi-Domain ScenariosAelim Ahn, Sooyeon Lee, Hyosun Wang, Chiwan Park et al.ICLR 2026
- CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World UncertaintyJohannes Kirmayr, Lukas Stappen, Elisabeth AndréACL 2026 · 5 citations
- Learning to Ask: When LLM Agents Meet Unclear InstructionWenxuan Wang, Juluan Shi, Zixuan Ling, Yuk-Kit Chan et al.EMNLP 2025 · 1 citation
