am-ELO: A Stable Framework for Arena-based LLM Evaluation
Zirui Liu, Jiatong Li, Yan Zhuang, Qi Liu, Shuanghong Shen, Jie Ouyang, Mingyue Cheng, Shijin Wang
摘要
Arena-based evaluation is a fundamental yet significant evaluation paradigm for modern AI models, especially large language models (LLMs). Existing framework based on ELO rating system suffers from the inevitable instability problem due to ranking inconsistency and the lack of attention to the varying abilities of annotators. In this paper, we introduce a novel stable arena framework to address these issues by enhancing the ELO Rating System. Specifically, we replace the iterative update method with a Maximum Likelihood Estimation (MLE) approach, m-ELO, and provide theoretical proof of the consistency and stability of the MLE approach for model ranking. Additionally, we proposed the am-ELO, which modify the Elo Rating's probability function to incorporate annotator abilities, enabling the simultaneous estimation of model scores and annotator reliability. Experiments demonstrate that this method ensures stability, proving that this framework offers a more robust, accurate, and stable evaluation method for LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Robust AI Evaluation through Maximal LotteriesHadi Khalaf, Serena Wang, Daniel Halpern, Itai Shapira 等ICML 2026 · 被引用 2 次
- Fewer Battles, More Gain: An Information-Efficient Framework for Arena-based LLM EvaluationZirui Liu, Xianquan Wang, Yan Zhuang, Jiatong Li 等ICLR 2026
- Rethinking Pretraining Data Detection for LLMs: From Local to GlobalChenye Ke, Yan Zhuang, Zirui Liu, Qi LiuICML 2026
它引用的顶会 Paper10
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning OptimizationYidong Wang, Zhuohao Yu, Wenjin Yao, Zhengran Zeng 等ICLR 2024 · 被引用 368 次
- tinyBenchmarks: evaluating LLMs with fewer examplesFelipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun 等ICML 2024 · 被引用 212 次
- Generative Judge for Evaluating AlignmentJunlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan 等ICLR 2024 · 被引用 173 次
相关 Paper
- Ranking Unraveled: Recipes for LLM Rankings in Head-to-Head AI CombatRoland Daynauth, Christopher Clarke, Krisztián Flautner, Lingjia Tang 等ACL 2025
- Elo Uncovered: Robustness and Best Practices in Language Model EvaluationMeriem Boubdir, Edward Kim, Beyza Ermis, Sara Hooker 等NeurIPS 2024 · 被引用 94 次
- Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language ModelsYanbin Yin, Kun Zhou, Zhen Wang, Xiangdong Zhang 等ACL 2026 · 被引用 2 次
- Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee DiscussionsRuochen Zhao, Wenxuan Zhang, Yew Ken Chia, Weiwen Xu 等ACL 2025 · 被引用 34 次
- GuessArena: Guess Who I Am? A Self-Adaptive Framework for Evaluating LLMs in Domain-Specific Knowledge and ReasoningQingchen Yu, Zifan Zheng, Ding Chen, Simin Niu 等ACL 2025 · 被引用 5 次
