Re-evaluating Open-ended Evaluation of Large Language Models
Siqi Liu, Ian Gemp, Luke Marris, Georgios Piliouras, Nicolas Heess, Marc Lanctot
Abstract
Evaluation has traditionally focused on ranking candidates for a specific skill. Modern generalist models, such as Large Language Models (LLMs), decidedly outpace this paradigm. Open-ended evaluation systems, where candidate models are compared on user-submitted prompts, have emerged as a popular solution. Despite their many advantages, we show that the current Elo-based rating systems can be susceptible to and even reinforce biases in data, intentional or accidental, due to their sensitivity to redundancies. To address this issue, we propose evaluation as a 3-player game, and introduce novel game-theoretic solution concepts to ensure robustness to redundancy. We show that our method leads to intuitive ratings and provide insights into the competitive landscape of LLM development.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- SpatiaLab: Can Vision-Language Models Perform Spatial Reasoning in the Wild?Azmine Toushik Wasi, Wahid Faisal, Abdur Rahman, Mahfuz Ahmed Anik et al.ICLR 2026 · 13 citations
- Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative SignalsZihan Dong, Zhixian Zhang, Yang Zhou, Can Jin et al.ICML 2026 · 2 citations
- UDA: Unsupervised Debiasing Alignment for Pair-wise LLM-as-a-JudgeYang Zhang, Cunxiang Wang, Lindong Wu, Wenbo Yu et al.AAAI 2026
- Revisiting Regularized Policy Optimization for Stable and Efficient Reinforcement Learning in Two-Player GamesKazuki Ota, Takayuki Osa, Motoki Omura, Tatsuya HaradaICML 2026
Related papers
- Ranking Unraveled: Recipes for LLM Rankings in Head-to-Head AI CombatRoland Daynauth, Christopher Clarke, Krisztián Flautner, Lingjia Tang et al.ACL 2025
- Elo Uncovered: Robustness and Best Practices in Language Model EvaluationMeriem Boubdir, Edward Kim, Beyza Ermis, Sara Hooker et al.NeurIPS 2024 · 94 citations
- Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy CompetitionKehua Feng, Keyan Ding, Hongzhi Tan, Kede Ma et al.ACL 2025
- Rethinking Generative Large Language Model Evaluation for Semantic ComprehensionFangyun Wei, Xi Chen, Lin LuoICML 2024 · 15 citations
- Beyond Utility: Evaluating LLM as RecommenderChumeng Jiang, Jiayin Wang, Weizhi Ma, Charles L. A. Clarke et al.WWW 2025 · 22 citations
