Ranking Unraveled: Recipes for LLM Rankings in Head-to-Head AI Combat
Roland Daynauth, Christopher Clarke, Krisztián Flautner, Lingjia Tang, Jason Mars
摘要
Evaluating large language models (LLMs) is a complex task. Pairwise ranking, where humans compare LLM outputs based on predefined criteria, has become a leading approach. By aggregating these comparisons through algorithms such as Elo, rankings across multiple LLMs can be derived. However, applying ranking algorithms in LLM evaluation presents several challenges. Traditional systems like Elo, designed initially for structured competitions such as chess, often produce inconsistent and unstable rankings due to the dynamic and context-dependent nature of LLM performance. Despite the increasing reliance on these methods, a systematic study of ranking algorithms for LLM evaluation remains lacking. This paper examines the effectiveness of various ranking systems for head-to-head LLM comparisons. We define key principles for robust ranking, conduct extensive evaluations of different ranking algorithms, and analyze their stability, accuracy, and sensitivity to real-world conditions. Our findings offer insights into the limitations of existing approaches and provide guidelines for selecting the most appropriate ranking method based on evaluation objectives and resource constraints.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang 等NeurIPS 2023 · 被引用 948 次
- Elo Uncovered: Robustness and Best Practices in Language Model EvaluationMeriem Boubdir, Edward Kim, Beyza Ermis, Sara Hooker 等NeurIPS 2024 · 被引用 94 次
相关 Paper
- Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy CompetitionKehua Feng, Keyan Ding, Hongzhi Tan, Kede Ma 等ACL 2025
- Rethinking Generative Large Language Model Evaluation for Semantic ComprehensionFangyun Wei, Xi Chen, Lin LuoICML 2024 · 被引用 15 次
- Re-evaluating Open-ended Evaluation of Large Language ModelsSiqi Liu, Ian Gemp, Luke Marris, Georgios Piliouras 等ICLR 2025
- am-ELO: A Stable Framework for Arena-based LLM EvaluationZirui Liu, Jiatong Li, Yan Zhuang, Qi Liu 等ICML 2025
- A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and RecommendationsMd. Tahmid Rahman Laskar, Sawsan Alqahtani, M. Saiful Bari, Mizanur Rahman 等EMNLP 2024 · 被引用 47 次
