ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox
Yuanyang Li, Xue Yang, Longyue Wang, Weihua Luo, Hongyang Chen
摘要
Current LLM agents are proficient at calling isolated APIs but struggle with the "last mile" of commercial software automation. In real-world scenarios, tools are not independent; they are atomic, interdependent, and prone to environmental noise. We introduce , a benchmark designed to evaluate agents in these rigorous conditions. Built on the Model Context Protocol (MCP), provides over 300 meticulously tested tools derived from 7 stateful sandboxes, ranging from office suites to financial systems. Unlike existing datasets, our benchmark utilizes a seed-driven architecture to simulate dynamic environment states and unpredictable API failures, ensuring a deterministic yet diverse evaluation. We evaluate various LLMs across full-context and RAG paradigms, revealing a stark performance gap: even top-tier models fail to exceed a 60% success rate, far trailing human performance 90%. Granular trajectory analysis identifies three fundamental bottlenecks: (1) as action spaces scale; (2) , where agents skip essential environment verifications; and (3) , a tendency to rationalize failure rather than pursuing recovery. These findings underscore the insufficiency of current agents for interdependent workflows, positioning as a critical testbed for the next generation of resilient autonomous systems. The codebase and benchmark implementation are publicly available at https://github.com/AIDC-AI/complex-mcp.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP ServersZhenting Wang, Qi Chang, Hemani Patel, Shashank Biju 等ICLR 2026 · 被引用 109 次
- TRAJECT-Bench: A Trajectory-Aware Benchmark for Evaluating Agentic Tool UsePengfei He, Zhenwei Dai, Bing He, Hui Liu 等ICLR 2026 · 被引用 46 次
- Search-o1: Agentic Search-Enhanced Large Reasoning ModelsXiaoxi Li, Guanting Dong, Jiajie Jin, Yuyao Zhang 等EMNLP 2025 · 被引用 12 次
- ChatDev: Communicative Agents for Software DevelopmentChen Qian, Wei Liu, Hongzhang Liu, Nuo Chen 等ACL 2024
- The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language ModelsShishir G. Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji 等ICML 2025
相关 Paper
- MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP ServersXuanjun Zong, Zhiqi Shen, Lei Wang, Yunshi Lan 等ICLR 2026 · 被引用 34 次
- MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP UseZijian Wu, Xiangyan Liu, Xinyuan Zhang, Lingjun Chen 等ICLR 2026 · 被引用 40 次
- MCP-Flow: Facilitating LLM Agents to Master Real-World, Diverse and Scaling MCP ToolsWenhao Wang, Peizhi Niu, Zhao Xu, Zhaoyu Chen 等ACL 2026 · 被引用 8 次
- MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM AgentsDongsen Zhang, Zekun Li, Xu Luo, Xuannan Liu 等ICLR 2026 · 被引用 47 次
- MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated ToolsZikang Guo, Benfeng Xu, Chiwei Zhu, Wentao Hong 等AAAI 2026 · 被引用 19 次
