Elo Uncovered: Robustness and Best Practices in Language Model Evaluation
Meriem Boubdir, Edward Kim, Beyza Ermis, Sara Hooker, Marzieh Fadaee
Abstract
In Natural Language Processing (NLP), the Elo rating system, originally designed for ranking players in dynamic games such as chess, is increasingly being used to evaluate Large Language Models (LLMs) through "A vs B" paired comparisons. However, while popular, the system's suitability for assessing entities with constant skill levels, such as LLMs, remains relatively unexplored. We study two fundamental axioms that evaluation methods should adhere to: reliability and transitivity. We conduct an extensive evaluation of Elo behavior across simulated and real-world scenarios, demonstrating that individual Elo computations can exhibit significant volatility. We show that both axioms are not always satisfied, raising questions about the reliability of current comparative evaluations of LLMs. If the current use of Elo scores is intended to substitute the costly head-to-head comparison of LLMs, it is crucial to ensure the ranking is as robust as possible. Guided by the axioms, our findings offer concrete guidelines for enhancing the reliability of LLM evaluation methods, suggesting a need for reassessment of existing comparative approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 37b3875d-2577-4397-8f28-fb27e145f8b9Cited by top-tier papers25
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- TabDPT: Scaling Tabular Foundation Models on Real DataJunwei Ma, Valentin Thomas, Rasa Hosseinzadeh, Alex Labach et al.NeurIPS 2025 · 118 citations
- Evaluating Language Model Agency Through NegotiationsTim R. Davidson, Veniamin Veselovsky, Michal Kosinski, Robert WestICLR 2024 · 52 citations
- A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and RecommendationsMd. Tahmid Rahman Laskar, Sawsan Alqahtani, M. Saiful Bari, Mizanur Rahman et al.EMNLP 2024 · 47 citations
- Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee DiscussionsRuochen Zhao, Wenxuan Zhang, Yew Ken Chia, Weiwen Xu et al.ACL 2025 · 34 citations
Builds on5
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Enhancing Chat Language Models by Scaling High-quality Instructional ConversationsNing Ding, Yulin Chen, Bokai Xu, Yujia Qin et al.EMNLP 2023 · 95 citations
- Nondeterminism and Instability in Neural Network OptimizationCecilia Summers, Michael J. DinneenICML 2021 · 55 citations
- Elo-MMR: A Rating System for Massive Multiplayer CompetitionsAram Ebtekar, Paul LiuWWW 2021 · 21 citations
- On the Challenges of Using Black-Box APIs for Toxicity Evaluation in ResearchLuiza Pozzobon, Beyza Ermis, Patrick Lewis, Sara HookerEMNLP 2023 · 19 citations
Related papers
- Ranking Unraveled: Recipes for LLM Rankings in Head-to-Head AI CombatRoland Daynauth, Christopher Clarke, Krisztián Flautner, Lingjia Tang et al.ACL 2025
- Re-evaluating Open-ended Evaluation of Large Language ModelsSiqi Liu, Ian Gemp, Luke Marris, Georgios Piliouras et al.ICLR 2025
- am-ELO: A Stable Framework for Arena-based LLM EvaluationZirui Liu, Jiatong Li, Yan Zhuang, Qi Liu et al.ICML 2025
- Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy CompetitionKehua Feng, Keyan Ding, Hongzhi Tan, Kede Ma et al.ACL 2025
- Investigating Non-Transitivity in LLM-as-a-JudgeYi Xu, Laura Ruis, Tim Rocktäschel, Robert KirkICML 2025
