SGVEF-LOOP: Coverage-Guided Progressive Topological Exploration and Fact-Grounded Metamorphic Evaluation for MCP Agents
Zenghao Liu, Yansong Zhang
Abstract
The rapid expansion of the Model Context Protocol (MCP) ecosystem introduces a combinatorially complex tool space, rendering existing frameworks inadequate for comprehensive agent evaluation. To address this problem, we propose SGVEF-LOOP, a coverage-guided framework for progressive topological exploration and fact-augmented metamorphic testing. SGVEF operates via a synergistic closed loop: it navigates sparse regions using adaptive sampling, synthesizes oracle-free metamorphic pairs grounded in static knowledge, enforces dual-constraint validation to ensure consistency and solvability, and leverages execution feedback to iteratively optimize exploration. Deploying this framework yields a high-fidelity benchmark achieving 100% node coverage and saturating 80.54% of the theoretical transition bound (estimated via Chao1). Evaluation of 8 diverse MCP Agents reveals capability stratification and exposes critical behavioral anomalies-such as reasoning instability-that conventional metrics fail to capture. Consequently, this work establishes a generalizable paradigm for scalable, rigorous agent evaluation in dynamic environments. Our code and dataset are available at https: //github.com/zh6zh/SGVEF-LOOP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b8e6d77a-d21b-402f-8812-9e0025de5114Builds on11
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu et al.ICLR 2024 · 1,469 citations
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu et al.ICLR 2024 · 748 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP ServersZhenting Wang, Qi Chang, Hemani Patel, Shashank Biju et al.ICLR 2026 · 109 citations
- MultiAgentBench : Evaluating the Collaboration and Competition of LLM agentsKunlun Zhu, Hongyi Du, Zhaochen Hong, Xiaocheng Yang et al.ACL 2025 · 97 citations
Related papers
- MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated ToolsZikang Guo, Benfeng Xu, Chiwei Zhu, Wentao Hong et al.AAAI 2026 · 19 citations
- DREAM: Deep Research Evaluation with Agentic MetricsElad Ben-Avraham, Changhao Li, Ron Dorfman, Roy Ganz et al.ACL 2026 · 2 citations
- TRACE: Trajectory-Aware Comprehensive Evaluation for Deep Research AgentsYanyu Chen, Jiyue Jiang, Jiahong Liu, Yifei Zhang et al.WWW 2026 · 4 citations
- OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use AgentsHongrui Jia, Jitong Liao, Xi Zhang, Haiyang Xu et al.ICLR 2026 · 24 citations
- SAGE-NAS: Synergizing LLM-Based Semantic Agent with Graph-Based Evaluator for Neural Architecture SearchKaiqi Lin, Jianping LuoICML 2026
