Pluralistic Leaderboards
Nika Haghtalab, Ariel Procaccia, Han Shao, Serena Wang, Kunhe Yang
Abstract
Recent leaderboard-based evaluations of large language models aggregate user feedback by fitting a Bradley--Terry model to pairwise comparisons, producing a single global ranking based on a latent quality score. While appealing for its simplicity, this approach is incompatible with heterogeneous preferences: when LLMs are used across diverse tasks and use cases, users who favor fundamentally different model behaviors can be systematically misrepresented when collapsed into a single quality score. To address this issue, we study pluralistic leaderboards that aim to remain stable with respect to heterogeneous user populations. Drawing on ideas from social choice theory, we adapt the notion of local stability, which requires that no model outside the top- positions is collectively preferred to the top- set by more than fraction of users. Building on techniques from the social choice literature, we design an alternative leaderboard mechanism that satisfies local stability while eliciting only pairwise comparisons per user, where is the size of the prefix for which stability is guaranteed. Using data from LMArena, we show that standard Bradley--Terry aggregation can violate local stability in practice, whereas our method provides substantially stronger stability guarantees.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2441bb4f-a2e7-4fd1-8bc5-fc2ec67757ddBuilds on10
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- Proportional Participatory Budgeting with Additive UtilitiesDominik Peters, Grzegorz Pierczynski, Piotr SkowronNeurIPS 2021 · 168 citations
- Representation with Incomplete VotesDaniel Halpern, Gregory Kehne, Ariel D. Procaccia, Jamie Tucker-Foltz et al.AAAI 2023 · 30 citations
- Approximately stable committee selectionZhihao Jiang, Kamesh Munagala, Kangning WangSTOC 2020 · 29 citations
Related papers
- Robust AI Evaluation through Maximal LotteriesHadi Khalaf, Serena Wang, Daniel Halpern, Itai Shapira et al.ICML 2026 · 2 citations
- Prompt-to-Leaderboard: Prompt-Adaptive LLM EvaluationsEvan Frick, Connor Chen, Joseph Tennyson, Tianle Li et al.ICML 2025
- Nonparametric LLM Evaluation from Preference DataDennis Frauen, Athiya Deviyani, Mihaela van der Schaar, Stefan FeuerriegelICML 2026 · 5 citations
- When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model LeaderboardsNorah A. Alzahrani, Hisham Abdullah Alyahya, Yazeed Alnumay, Sultan Alrashed et al.ACL 2024 · 13 citations
- A Judge-Aware Ranking Framework for Evaluating Large Language Models without Ground TruthMingyuan Xu, Xinzi Tan, Jiawei Wu, Doudou ZhouICML 2026
