Re-evaluating Open-ended Evaluation of Large Language Models
Siqi Liu, Ian Gemp, Luke Marris, Georgios Piliouras, Nicolas Heess, Marc Lanctot
摘要
Evaluation has traditionally focused on ranking candidates for a specific skill. Modern generalist models, such as Large Language Models (LLMs), decidedly outpace this paradigm. Open-ended evaluation systems, where candidate models are compared on user-submitted prompts, have emerged as a popular solution. Despite their many advantages, we show that the current Elo-based rating systems can be susceptible to and even reinforce biases in data, intentional or accidental, due to their sensitivity to redundancies. To address this issue, we propose evaluation as a 3-player game, and introduce novel game-theoretic solution concepts to ensure robustness to redundancy. We show that our method leads to intuitive ratings and provide insights into the competitive landscape of LLM development.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- SpatiaLab: Can Vision-Language Models Perform Spatial Reasoning in the Wild?Azmine Toushik Wasi, Wahid Faisal, Abdur Rahman, Mahfuz Ahmed Anik 等ICLR 2026 · 被引用 13 次
- Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative SignalsZihan Dong, Zhixian Zhang, Yang Zhou, Can Jin 等ICML 2026 · 被引用 2 次
- UDA: Unsupervised Debiasing Alignment for Pair-wise LLM-as-a-JudgeYang Zhang, Cunxiang Wang, Lindong Wu, Wenbo Yu 等AAAI 2026
- Revisiting Regularized Policy Optimization for Stable and Efficient Reinforcement Learning in Two-Player GamesKazuki Ota, Takayuki Osa, Motoki Omura, Tatsuya HaradaICML 2026
相关 Paper
- Ranking Unraveled: Recipes for LLM Rankings in Head-to-Head AI CombatRoland Daynauth, Christopher Clarke, Krisztián Flautner, Lingjia Tang 等ACL 2025
- Elo Uncovered: Robustness and Best Practices in Language Model EvaluationMeriem Boubdir, Edward Kim, Beyza Ermis, Sara Hooker 等NeurIPS 2024 · 被引用 94 次
- Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy CompetitionKehua Feng, Keyan Ding, Hongzhi Tan, Kede Ma 等ACL 2025
- Rethinking Generative Large Language Model Evaluation for Semantic ComprehensionFangyun Wei, Xi Chen, Lin LuoICML 2024 · 被引用 15 次
- Beyond Utility: Evaluating LLM as RecommenderChumeng Jiang, Jiayin Wang, Weizhi Ma, Charles L. A. Clarke 等WWW 2025 · 被引用 22 次
