MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
Zikang Guo, Benfeng Xu, Chiwei Zhu, Wentao Hong, Xiaorui Wang, Zhendong Mao
Abstract
The Model Context Protocol (MCP) is rapidly emerging as a pivotal open standard, designed to enhance agent-tool integration and interoperability, and is positioned to unlock a new era of powerful, interconnected, and genuinely utilitarian agentic AI. However, despite MCP's growing adoption, existing benchmarks often fail to capture real-world agent performance within this new paradigm, leading to a distorted perception of their true operational value and an inability to reliably differentiate proficiencies. To bridge this critical evaluation gap, we introduce MCP-AgentBench—a comprehensive benchmark specifically engineered to rigorously assess language agent capabilities in MCP-mediated tool interactions. Core contributions of MCP-AgentBench include: the establishment of a robust MCP testbed comprising 33 operational servers with 188 distinct tools; the development of a benchmark featuring 600 systematically designed queries distributed across 6 distinct categories of varying interaction complexity; and the introduction of MCP-Eval, a novel outcome-oriented evaluation methodology prioritizing real-world task success. Through extensive empirical evaluation of leading language agents, we provide foundational insights. MCP-AgentBench aims to equip the research community with a standardized and reliable framework to build, validate, and advance agents capable of fully leveraging MCP's transformative benefits, thereby accelerating progress toward truly capable and interoperable AI systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext de8687fc-82b5-403b-b511-28597dbbffa8Cited by top-tier papers5
- The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task ExecutionJunlong Li, Wenshuo Zhao, Jian Zhao, Weihao Zeng et al.ICLR 2026 · 73 citations
- MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP UseZijian Wu, Xiangyan Liu, Xinyuan Zhang, Lingjun Chen et al.ICLR 2026 · 40 citations
- FS-Researcher: Test-Time Scaling for Long-Horizon Research Tasks with File-System-Based AgentsChiwei Zhu, Benfeng Xu, Mingxuan Du, Shaohan Wang et al.ACL 2026 · 2 citations
- MedMCP-Calc: Benchmarking LLMs for Realistic Medical Calculator Scenarios via MCP IntegrationYakun Zhu, Yutong Huang, Shengqian Qin, Zhongzhen Huang et al.ACL 2026 · 2 citations
- AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World EnvironmentsZhiheng Xi, Dingwen Yang, Jiaqi Liu, Jixuan Huang et al.ACL 2026
Builds on4
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 1,715 citations
- API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMsMinghao Li, Yingxiu Zhao, Bowen Yu, Feifan Song et al.EMNLP 2023 · 72 citations
- MetaGPT: Meta Programming for A Multi-Agent Collaborative FrameworkSirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng et al.ICLR 2024
Related papers
- MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP ServersXuanjun Zong, Zhiqi Shen, Lei Wang, Yunshi Lan et al.ICLR 2026 · 34 citations
- MCP-Persona: Benchmarking LLM Agents on Real-World Personal Applications via Environment SimulationWenhao Wang, Peizhi Niu, Gongyi Zou, Xiyuan Yang et al.ICML 2026 · 1 citation
- MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM AgentsDongsen Zhang, Zekun Li, Xu Luo, Xuannan Liu et al.ICLR 2026 · 47 citations
- MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP ServersZhenting Wang, Qi Chang, Hemani Patel, Shashank Biju et al.ICLR 2026 · 109 citations
- OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use AgentsHongrui Jia, Jitong Liao, Xi Zhang, Haiyang Xu et al.ICLR 2026 · 24 citations
