MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Zhenting Wang, Qi Chang, Hemani Patel, Shashank Biju, Cheng-En Wu, Quan Liu, Aolin Ding, Alireza Rezazadeh, Ankit Shah, Yujia Bao, Eugene Siow
Abstract
We introduce M C P-Be nc h, a benchmark for evaluating large language models (LLMs) on realistic, multi-step tasks that demand tool use, cross-tool coordination, precise parameter control, and planning/reasoning for solving tasks. Built on the Model Context Protocol (MCP), M C P-Be nc h connects LLMs to 28 representative live MCP servers spanning 250 tools across domains such as finance, traveling, scientific computing, and academic search. Unlike prior API-based benchmarks, each MCP server provides a set of complementary tools designed to work together, enabling the construction of authentic, multi-step tasks with rich input-output coupling. Also, tasks in M C P-Be nc h test agents' ability to retrieve relevant tools from fuzzy instructions without explicit tool names, plan multi-hop execution trajectories for complex objectives, ground responses in intermediate tool outputs, and orchestrate cross-domain workflows-capabilities not adequately evaluated by existing benchmarks that rely on explicit tool specifications, shallow few-step workflows, and isolated domain operations. We propose a multi-faceted evaluation framework covering tool-level schema understanding and usage, trajectorylevel planning and task completion. Experiments on 20 advanced LLMs reveal persistent challenges in M C P-Be nc h. Code and data: https://github.com/Accenture/mcp-bench .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 115c1e46-9da6-4d83-aefe-d5a0faea7f04Cited by top-tier papers11
- Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous EnvironmentsRomain Froger, Pierre Andrews, Matteo Bettini, Amar Budhiraja et al.ICLR 2026 · 29 citations
- Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing AgentsMuyu He, Anand Kumar, Soumyadeep Bakshi, James Zou et al.ACL 2026 · 8 citations
- InfoMosaic-Bench: Evaluating Multi-Source Information Seeking in Tool-Augmented AgentsYaxin Du, Yuanshuo Zhang, Xiyuan Yang, Yifan Zhou et al.ICLR 2026 · 3 citations
- DV-World: Benchmarking Data Visualization Agents in Real-World ScenariosJinxiang Meng, Shaoping Huang, Fangyu Lei, Jingyu Guo et al.ICML 2026 · 1 citation
- ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool SandboxYuanyang Li, Xue Yang, Longyue Wang, Weihua Luo et al.ICML 2026 · 1 citation
Builds on9
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu et al.ICLR 2024 · 1,469 citations
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu et al.ICLR 2024 · 748 citations
- TravelPlanner: A Benchmark for Real-World Planning with Language AgentsJian Xie, Kai Zhang, Jiangjie Chen, Tinghui Zhu et al.ICML 2024 · 376 citations
Related papers
- MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP ServersXuanjun Zong, Zhiqi Shen, Lei Wang, Yunshi Lan et al.ICLR 2026 · 34 citations
- MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM AgentsDongsen Zhang, Zekun Li, Xu Luo, Xuannan Liu et al.ICLR 2026 · 47 citations
- MCP-Persona: Benchmarking LLM Agents on Real-World Personal Applications via Environment SimulationWenhao Wang, Peizhi Niu, Gongyi Zou, Xiyuan Yang et al.ICML 2026 · 1 citation
- TPS-Bench: Evaluating AI Agents' Tool Planning & Scheduling Abilities in Compounding TasksHanwen Xu, Xuyao Huang, Yuzhe Liu, Zhijie DengACL 2026 · 2 citations
- MCP-Flow: Facilitating LLM Agents to Master Real-World, Diverse and Scaling MCP ToolsWenhao Wang, Peizhi Niu, Zhao Xu, Zhaoyu Chen et al.ACL 2026 · 8 citations
