Agent-as-a-Judge: Evaluate Agents with Agents
Mingchen Zhuge, Changsheng Zhao, Dylan R. Ashley, Wenyi Wang, Dmitrii Khizbullin, Yunyang Xiong, Zechun Liu, Ernie Chang, Raghuraman Krishnamoorthi, Yuandong Tian, Yangyang Shi, Vikas Chandra, Jürgen Schmidhuber
摘要
Contemporary evaluation techniques are inadequate for agentic systems. These approaches either focus exclusively on final outcomesignoring the step-by-step nature of the thinking done by agentic systems-or require excessive manual labour. To address this, we introduce the Agent-as-a-Judge framework, wherein agentic systems are used to evaluate agentic systems. This is a natural extension of the LLM-as-a-Judge framework, incorporating agentic features that enable intermediate feedback for the entire tasksolving processes for more precise evaluations. We apply the Agent-as-a-Judge framework to the task of code generation. To overcome issues with existing benchmarks and provide a proof-of-concept testbed for Agent-as-a-Judge, we present DevAI, a new benchmark of 55 realistic AI code generation tasks. DevAI includes rich manual annotations, like a total of 365 hierarchical solution requirements, which make it particularly suitable for an agentic evaluator. We benchmark three of the top code-generating agentic systems using Agent-as-a-Judge and find that our framework dramatically outperforms LLMas-a-Judge and is as reliable as our human evaluation baseline. Altogether, we believe that this work represents a concrete step towards enabling vastly more sophisticated agentic systems. To help that, our dataset and the full implementation of Agent-as-a-Judge will be publically available at https://github.com/metau to-ai/agent-as-a-judge First four authors made core contributions. KAUST crafted the dataset. Work done while Mingchen was interning at Meta, with Changsheng leading.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper30
- AgentAuditor: Human-level Safety and Security Evaluation for LLM AgentsHanjun Luo, Shenyu Dai, Chiming Ni, Xinfeng Li 等NeurIPS 2025 · 被引用 98 次
- The Limits of Inference Scaling Through ResamplingBenedikt Stroebl, Sayash Kapoor, Arvind NarayananICLR 2026 · 被引用 38 次
- Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human EvaluationJiaju Chen, Yuxuan Lu, Xiaojie Wang, Huimin Zeng 等ACL 2026 · 被引用 30 次
- DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent SystemsMing Ma, Jue Zhang, Fangkai Yang, Yu Kang 等ICLR 2026 · 被引用 24 次
- Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement LearningRan Xu, Jingjing Chen, Jiayu Ye, Yu Wu 等ICLR 2026 · 被引用 17 次
它引用的顶会 Paper19
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret 等NeurIPS 2024 · 被引用 2,059 次
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin 等NeurIPS 2023 · 被引用 1,975 次
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu 等ICLR 2024 · 被引用 748 次
- DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationYuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang 等ICML 2023 · 被引用 504 次
相关 Paper
- Interactive Evaluation of Large Language Models for Multi-Requirement Software Engineering TasksDimitrios Rontogiannis, Maxime Peyrard, Nicolas Mario Baldwin, Martin Josifoski 等AAAI 2026 · 被引用 1 次
- AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World ContextsKeyu Li, Junhao Shi, Yang Xiao, Mohan Jiang 等ACL 2026 · 被引用 14 次
- WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development QualityChunyang Li, Yilun Zheng, Xinting Huang, Tianqing Fang 等ICLR 2026 · 被引用 14 次
- DA-Code: Agent Data Science Code Generation Benchmark for Large Language ModelsYiming Huang, Jianwen Luo, Yan Yu, Yitong Zhang 等EMNLP 2024 · 被引用 7 次
- LegalAgentBench: Evaluating LLM Agents in Legal DomainHaitao Li, Junjie Chen, Jingli Yang, Qingyao Ai 等ACL 2025
