Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions
Ruochen Zhao, Wenxuan Zhang, Yew Ken Chia, Weiwen Xu, Deli Zhao, Lidong Bing
Abstract
As LLMs continuously evolve, there is an urgent need for a reliable evaluation method that delivers trustworthy results promptly. Currently, static benchmarks suffer from inflexibility and unreliability, leading users to prefer human voting platforms like Chatbot Arena. However, human evaluations require significant manual effort. To address this, we propose the Auto-Arena, an innovative framework that automates the entire evaluation process using LLM-powered agents. Firstly, an LLM examiner generates questions. Then, two LLM candidates engage in a multi-round peer battle based on individual questions, aiming at revealing their true performance differences. Finally, a committee of LLM judges collaboratively discusses and decides the winner, reducing bias and enhancing fairness. During the peer battles, we observe intriguing scenarios where the LLM candidates display competitive behaviors and even learn from the opponents. In our extensive experiments involving 15 recent LLMs, Auto-Arena shows a 92.14% correlation with human preferences, surpassing all previous expert-annotated benchmarks without any manual efforts. As a result, Auto-Arena offers a promising alternative to current human evaluation platforms for evaluating LLMs automatically. * Work done while the author was an intern at DAMO Academy, Alibaba Group. † Wenxuan Zhang is the corresponding author. ‡ Yew Ken Chia is under the Joint Ph.D. Program between DAMO Academy and SUTD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0af12314-0cb7-4315-a943-76ccdcf85b57Cited by top-tier papers8
- One Battle After Another: Probing LLMs' Limits on Multi-Turn Instruction Following with a Benchmark Evolving FrameworkQi Jia, Ye Shen, Xiujie Song, Kaiwei Zhang et al.ACL 2026 · 3 citations
- DR-Arena: an Automated Evaluation Framework for Deep Research AgentsYiwen Gao, Ruochen Zhao, Yang Deng, Wenxuan ZhangACL 2026 · 2 citations
- EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVerPukun Zhao, Longxiang Wang, Miaowei Wang, Chen Chen et al.AAAI 2026 · 2 citations
- Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence AttributionXiaoou Liu, Tiejin Chen, Dengjia Zhang, Yaqing Wang et al.ICML 2026 · 1 citation
- SCAN: Structured Capability Assessment and Navigation for LLMsZongqi Wang, Tianle Gu, Chen Gong, Xin Tian et al.ACL 2026 · 1 citation
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang et al.NeurIPS 2023 · 948 citations
Related papers
- Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language ModelsYanbin Yin, Kun Zhou, Zhen Wang, Xiangdong Zhang et al.ACL 2026 · 2 citations
- VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User SimulationZiyang Luo, Haoning Wu, Dongxu Li, Jing Ma et al.CVPR 2025
- MixEval: Deriving Wisdom of the Crowd from LLM Benchmark MixturesJinjie Ni, Fuzhao Xue, Xiang Yue, Yuntian Deng et al.NeurIPS 2024 · 88 citations
- From Crowdsourced Data to High-quality Benchmarks: Arena-Hard and Benchbuilder PipelineTianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap et al.ICML 2025
- LitReview Arena: Evaluating Literature Review Agents with Battle-style Peer Review PlatformRuotong Zhao, Zhiyu Chen, Xurui Liu, Haidong Xue et al.ICML 2026
