ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox
Yuanyang Li, Xue Yang, Longyue Wang, Weihua Luo, Hongyang Chen
Abstract
Current LLM agents are proficient at calling isolated APIs but struggle with the "last mile" of commercial software automation. In real-world scenarios, tools are not independent; they are atomic, interdependent, and prone to environmental noise. We introduce , a benchmark designed to evaluate agents in these rigorous conditions. Built on the Model Context Protocol (MCP), provides over 300 meticulously tested tools derived from 7 stateful sandboxes, ranging from office suites to financial systems. Unlike existing datasets, our benchmark utilizes a seed-driven architecture to simulate dynamic environment states and unpredictable API failures, ensuring a deterministic yet diverse evaluation. We evaluate various LLMs across full-context and RAG paradigms, revealing a stark performance gap: even top-tier models fail to exceed a 60% success rate, far trailing human performance 90%. Granular trajectory analysis identifies three fundamental bottlenecks: (1) as action spaces scale; (2) , where agents skip essential environment verifications; and (3) , a tendency to rationalize failure rather than pursuing recovery. These findings underscore the insufficiency of current agents for interdependent workflows, positioning as a critical testbed for the next generation of resilient autonomous systems. The codebase and benchmark implementation are publicly available at https://github.com/AIDC-AI/complex-mcp.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP ServersZhenting Wang, Qi Chang, Hemani Patel, Shashank Biju et al.ICLR 2026 · 109 citations
- TRAJECT-Bench: A Trajectory-Aware Benchmark for Evaluating Agentic Tool UsePengfei He, Zhenwei Dai, Bing He, Hui Liu et al.ICLR 2026 · 46 citations
- Search-o1: Agentic Search-Enhanced Large Reasoning ModelsXiaoxi Li, Guanting Dong, Jiajie Jin, Yuyao Zhang et al.EMNLP 2025 · 12 citations
- ChatDev: Communicative Agents for Software DevelopmentChen Qian, Wei Liu, Hongzhang Liu, Nuo Chen et al.ACL 2024
- The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language ModelsShishir G. Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji et al.ICML 2025
Related papers
- MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP ServersXuanjun Zong, Zhiqi Shen, Lei Wang, Yunshi Lan et al.ICLR 2026 · 34 citations
- MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP UseZijian Wu, Xiangyan Liu, Xinyuan Zhang, Lingjun Chen et al.ICLR 2026 · 40 citations
- MCP-Flow: Facilitating LLM Agents to Master Real-World, Diverse and Scaling MCP ToolsWenhao Wang, Peizhi Niu, Zhao Xu, Zhaoyu Chen et al.ACL 2026 · 8 citations
- MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM AgentsDongsen Zhang, Zekun Li, Xu Luo, Xuannan Liu et al.ICLR 2026 · 47 citations
- MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated ToolsZikang Guo, Benfeng Xu, Chiwei Zhu, Wentao Hong et al.AAAI 2026 · 19 citations
