Fewer Battles, More Gain: An Information-Efficient Framework for Arena-based LLM Evaluation
Zirui Liu, Xianquan Wang, Yan Zhuang, Jiatong Li, Qi Liu, Shuanghong Shen, Mingyue Cheng, Shijin Wang
摘要
Arena-based evaluation has become a key method for assessing large language models (LLMs) through head-to-head model comparisons, closely reflecting human preferences. However, current arena rating systems (e.g., ELO rating system) often suffer from inefficiencies due to exhaustive or random model pair annotations, leading to redundant evaluations, longer evaluation times, and lower overall efficiency. To address these challenges, we propose a novel adaptive modelpair selection algorithm. By leveraging the asymptotic normality of LLM ability estimation under sparse conditions, our approach strategically selects highvalue model pairs, focusing on confrontations with the lowest variance. Specifically, we introduce Fisher information as a metric to guide model pair selection, optimizing the evaluation process through A-optimality and D-optimality. A-optimality minimizes estimation variance, ensuring balanced reliability across models, while D-optimality reduces uncertainty by maximizing the determinant of the Fisher Information Matrix. Extensive experiments on both simulated and real-world datasets demonstrate that our method outperforms existing approaches in terms of information efficiency and result reliability. Notably, our method offers a flexible, general toolkit that can be easily integrated into existing arena-based platforms, greatly improving scalability and efficiency for largescale LLM evaluations. Our code is publicly available to promote reproducibility at https://github.com/Liuz-rui/Adaptive-Arena .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Re³: Relevance & Recency Retrieval for Mitigating Temporal HallucinationJiawei Cao, Jie Ouyang, Mingyue Cheng, Zhaomeng Zhou 等ACL 2026
- Rethinking Pretraining Data Detection for LLMs: From Local to GlobalChenye Ke, Yan Zhuang, Zirui Liu, Qi LiuICML 2026
它引用的顶会 Paper19
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- Data Selection for Language Models via Importance ResamplingSang Michael Xie, Shibani Santurkar, Tengyu Ma, Percy LiangNeurIPS 2023 · 被引用 383 次
- tinyBenchmarks: evaluating LLMs with fewer examplesFelipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun 等ICML 2024 · 被引用 212 次
- LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language ModelsZhiqiang Hu, Lei Wang, Yihuai Lan, Wanyu Xu 等EMNLP 2023 · 被引用 200 次
相关 Paper
- am-ELO: A Stable Framework for Arena-based LLM EvaluationZirui Liu, Jiatong Li, Yan Zhuang, Qi Liu 等ICML 2025
- MixEval: Deriving Wisdom of the Crowd from LLM Benchmark MixturesJinjie Ni, Fuzhao Xue, Xiang Yue, Yuntian Deng 等NeurIPS 2024 · 被引用 88 次
- Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language ModelsYanbin Yin, Kun Zhou, Zhen Wang, Xiangdong Zhang 等ACL 2026 · 被引用 2 次
- GuessArena: Guess Who I Am? A Self-Adaptive Framework for Evaluating LLMs in Domain-Specific Knowledge and ReasoningQingchen Yu, Zifan Zheng, Ding Chen, Simin Niu 等ACL 2025 · 被引用 5 次
- Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy CompetitionKehua Feng, Keyan Ding, Hongzhi Tan, Kede Ma 等ACL 2025
