Pluralistic Leaderboards
Nika Haghtalab, Ariel Procaccia, Han Shao, Serena Wang, Kunhe Yang
摘要
Recent leaderboard-based evaluations of large language models aggregate user feedback by fitting a Bradley--Terry model to pairwise comparisons, producing a single global ranking based on a latent quality score. While appealing for its simplicity, this approach is incompatible with heterogeneous preferences: when LLMs are used across diverse tasks and use cases, users who favor fundamentally different model behaviors can be systematically misrepresented when collapsed into a single quality score. To address this issue, we study pluralistic leaderboards that aim to remain stable with respect to heterogeneous user populations. Drawing on ideas from social choice theory, we adapt the notion of local stability, which requires that no model outside the top- positions is collectively preferred to the top- set by more than fraction of users. Building on techniques from the social choice literature, we design an alternative leaderboard mechanism that satisfies local stability while eliciting only pairwise comparisons per user, where is the size of the prefix for which stability is guaranteed. Using data from LMArena, we show that standard Bradley--Terry aggregation can violate local stability in practice, whereas our method provides substantially stronger stability guarantees.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- Proportional Participatory Budgeting with Additive UtilitiesDominik Peters, Grzegorz Pierczynski, Piotr SkowronNeurIPS 2021 · 被引用 168 次
- Representation with Incomplete VotesDaniel Halpern, Gregory Kehne, Ariel D. Procaccia, Jamie Tucker-Foltz 等AAAI 2023 · 被引用 30 次
- Approximately stable committee selectionZhihao Jiang, Kamesh Munagala, Kangning WangSTOC 2020 · 被引用 29 次
相关 Paper
- Robust AI Evaluation through Maximal LotteriesHadi Khalaf, Serena Wang, Daniel Halpern, Itai Shapira 等ICML 2026 · 被引用 2 次
- Prompt-to-Leaderboard: Prompt-Adaptive LLM EvaluationsEvan Frick, Connor Chen, Joseph Tennyson, Tianle Li 等ICML 2025
- Nonparametric LLM Evaluation from Preference DataDennis Frauen, Athiya Deviyani, Mihaela van der Schaar, Stefan FeuerriegelICML 2026 · 被引用 5 次
- When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model LeaderboardsNorah A. Alzahrani, Hisham Abdullah Alyahya, Yazeed Alnumay, Sultan Alrashed 等ACL 2024 · 被引用 13 次
- A Judge-Aware Ranking Framework for Evaluating Large Language Models without Ground TruthMingyuan Xu, Xinzi Tan, Jiawei Wu, Doudou ZhouICML 2026
