Lune

ICML2026Top-tier venue

Pluralistic Leaderboards

Nika Haghtalab, Ariel Procaccia, Han Shao, Serena Wang, Kunhe Yang

2026Year

Abstract

Recent leaderboard-based evaluations of large language models aggregate user feedback by fitting a Bradley--Terry model to pairwise comparisons, producing a single global ranking based on a latent quality score. While appealing for its simplicity, this approach is incompatible with heterogeneous preferences: when LLMs are used across diverse tasks and use cases, users who favor fundamentally different model behaviors can be systematically misrepresented when collapsed into a single quality score. To address this issue, we study pluralistic leaderboards that aim to remain stable with respect to heterogeneous user populations. Drawing on ideas from social choice theory, we adapt the notion of local stability, which requires that no model outside the top-kk positions is collectively preferred to the top-kk set by more than O(1/k)O(1/k) fraction of users. Building on techniques from the social choice literature, we design an alternative leaderboard mechanism that satisfies local stability while eliciting only O~(k)\widetilde{O}(k) pairwise comparisons per user, where kk is the size of the prefix for which stability is guaranteed. Using data from LMArena, we show that standard Bradley--Terry aggregation can violate local stability in practice, whereas our method provides substantially stronger stability guarantees.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 2441bb4f-a2e7-4fd1-8bc5-fc2ec67757dd

Builds on10

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines