ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, Zhiyuan Liu
Abstract
Text evaluation has historically posed significant challenges, often demanding substantial labor and time cost. With the emergence of large language models (LLMs), researchers have explored LLMs' potential as alternatives for human evaluation. While these single-agent-based approaches show promise, experimental results suggest that further advancements are needed to bridge the gap between their current effectiveness and human-level evaluation quality. Recognizing that best practices of human evaluation processes often involve multiple human annotators collaborating in the evaluation, we resort to a multi-agent debate framework, moving beyond single-agent prompting strategies. The multi-agentbased approach enables a group of LLMs to synergize with an array of intelligent counterparts, harnessing their distinct capabilities and expertise to enhance efficiency and effectiveness in handling intricate tasks. In this paper, we construct a multi-agent referee team called ChatEval to autonomously discuss and evaluate the quality of generated responses from different models on open-ended questions and traditional natural language generation (NLG) tasks. We derive insights and lessons from practical scenarios where humans instigate group discussions for brainstorming and propose different communication strategies within ChatEval. Our experiments on two benchmark tasks illustrate that ChatEval delivers superior accuracy and correlation in alignment with human assessment. Furthermore, we find that the diverse role prompts (different personas) are essential in the multi-agent debate process; that is, utilizing the same role description in the prompt can lead to a degradation in performance. Our qualitative analysis also shows that ChatEval transcends mere textual scoring, offering a humanmimicking evaluation process for reliable assessments. Our code is available at https://github.com/chanchimin/ChatEval .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da8fb2ac-6b1f-44eb-894e-2e345f95b72cCited by top-tier papers171
- AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent BehaviorsWeize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang et al.ICLR 2024 · 594 citations
- The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context LearningBill Yuchen Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri et al.ICLR 2024 · 299 citations
- SOTOPIA: Interactive Evaluation for Social Intelligence in Language AgentsXuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang et al.ICLR 2024 · 288 citations
- Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent CollaborationJunyang Wang, Haiyang Xu, Haitao Jia, Xi Zhang et al.NeurIPS 2024 · 245 citations
- MAGIS: LLM-Based Multi-Agent Framework for GitHub Issue ResolutionWei Tao, Yucheng Zhou, Yanlin Wang, Wenqiang Zhang et al.NeurIPS 2024 · 210 citations
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
Related papers
- Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human EvaluationJiaju Chen, Yuxuan Lu, Xiaojie Wang, Huimin Zeng et al.ACL 2026 · 30 citations
- Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMsAndries P. Smit, Nathan Grinsztajn, Paul Duckworth, Thomas D. Barrett et al.ICML 2024 · 82 citations
- Auto-PRE: An Automatic and Cost-Efficient Peer-Review Framework for Language Generation EvaluationJunjie Chen, Weihang Su, Zhumin Chu, Haitao Li et al.AAAI 2026
- IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question AnsweringRuosen Li, Ruochen Li, Barry Wang, Xinya DuNeurIPS 2024 · 26 citations
- Multi-Agent Debate for LLM Judges with Adaptive Stability DetectionTianyu Hu, Zhen Tan, Song Wang, Huaizhi Qu et al.NeurIPS 2025 · 25 citations
