Arena-lite: Efficient and Reliable Large Language Model Evaluation via Tournament-Based Direct Comparisons
Seonil Son, Ju-Min Oh, Heegon Jin, Cheolhun Jang, Jeongbeom Jeong, Kuntae Kim
Abstract
As Large Language Models (LLMs) expand across domains, LLM judges have become essential for systems evaluation. Current benchmarks typically compare system outputs against baselines. This baseline-mediated approach, though convenient, yields lower reliability than direct comparison between systems. We propose Arena-Lite which integrates tournament structure on top of head-to-head comparison. The application of a tournament structure and direct comparison eliminates the need for baseline outputs, reduces the number of required comparisons, and allows higher reliability in system rankings. We conducted two experiments: (1) controlled stochastic modeling and (2) empirical validation with a real LLM judge. Those experiments collectively demonstrate that Arena-Lite consistently achieves higher reliability with fewer comparisons, even with smaller datasets or weaker judges. We release an easy-to-use web demonstration and code to foster adoption of Arena-Lite, streamlining model selection across research and industry communities. Arena-Lite demo and code are available on https://huggingface. co/spaces/NCSOFT/ArenaLite
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 805933ea-b8a2-4801-b6c7-6ec6c0b78ce1Cited by top-tier papers1
Ask how each one uses itBuilds on5
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- PAL: Program-aided Language ModelsLuyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon et al.ICML 2023 · 700 citations
- UniCBE: An Uniformity-driven Comparing Based Evaluation Framework with Unified Multi-Objective OptimizationPeiwen Yuan, Shaoxiong Feng, Yiwei Li, Xinglin Wang et al.ICLR 2025
- JuStRank: Benchmarking LLM Judges for System RankingAriel Gera, Odellia Boni, Yotam Perlitz, Roy Bar-Haim et al.ACL 2025
Related papers
- Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language ModelsYanbin Yin, Kun Zhou, Zhen Wang, Xiangdong Zhang et al.ACL 2026 · 2 citations
- LitReview Arena: Evaluating Literature Review Agents with Battle-style Peer Review PlatformRuotong Zhao, Zhiyu Chen, Xurui Liu, Haidong Xue et al.ICML 2026
- MixEval: Deriving Wisdom of the Crowd from LLM Benchmark MixturesJinjie Ni, Fuzhao Xue, Xiang Yue, Yuntian Deng et al.NeurIPS 2024 · 88 citations
- Fewer Battles, More Gain: An Information-Efficient Framework for Arena-based LLM EvaluationZirui Liu, Xianquan Wang, Yan Zhuang, Jiatong Li et al.ICLR 2026
- Ranking Unraveled: Recipes for LLM Rankings in Head-to-Head AI CombatRoland Daynauth, Christopher Clarke, Krisztián Flautner, Lingjia Tang et al.ACL 2025
