Inherent Trade-Offs between Diversity and Stability in Multi-Task Benchmarks
Guanhua Zhang, Moritz Hardt
Abstract
We examine multi-task benchmarks in machine learning through the lens of social choice theory. We draw an analogy between benchmarks and electoral systems, where models are candidates and tasks are voters. This suggests a distinction between cardinal and ordinal benchmark systems. The former aggregate numerical scores into one model ranking; the latter aggregate rankings for each task. We apply Arrow's impossibility theorem to ordinal benchmarks to highlight the inherent limitations of ordinal systems, particularly their sensitivity to the inclusion of irrelevant models. Inspired by Arrow's theorem, we empirically demonstrate a strong trade-off between diversity and sensitivity to irrelevant changes in existing multi-task benchmarks. Our result is based on new quantitative measures of diversity and sensitivity that we introduce. Sensitivity quantifies the impact that irrelevant changes to tasks have on a benchmark. Diversity captures the degree of disagreement in model rankings across tasks. We develop efficient approximation algorithms for both measures, as exact computation is computationally challenging. Through extensive experiments on seven cardinal benchmarks and eleven ordinal benchmarks, we demonstrate a clear trade-off between diversity and stability: The more diverse a multi-task benchmark, the more sensitive to trivial changes it is. Additionally, we show that the aggregated rankings of existing benchmarks are highly unstable under irrelevant changes. The codes and data are available at https://socialfoundations. github.io/benchbench/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ec7af4ac-0056-4971-8ad5-dc7a70c7a173Cited by top-tier papers8
- Train-before-Test Harmonizes Language Model RankingsGuanhua Zhang, Ricardo Dominguez-Olmedo, Moritz HardtICLR 2026 · 13 citations
- Multivariate Stochastic Dominance via Optimal Transport and Applications to Models BenchmarkingGabriel Rioux, Apoorva Nitsure, Mattia Rigotti, Kristjan H. Greenewald et al.NeurIPS 2024 · 6 citations
- Nonparametric LLM Evaluation from Preference DataDennis Frauen, Athiya Deviyani, Mihaela van der Schaar, Stefan FeuerriegelICML 2026 · 5 citations
- Metritocracy: Representative Metrics for Lite BenchmarksAriel D. Procaccia, Ben Schiffer, Serena Wang, Shirley ZhangNeurIPS 2025 · 2 citations
- Robust AI Evaluation through Maximal LotteriesHadi Khalaf, Serena Wang, Daniel Halpern, Itai Shapira et al.ICML 2026 · 2 citations
Builds on10
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- From ImageNet to Image Classification: Contextualizing Progress on BenchmarksDimitris Tsipras, Shibani Santurkar, Logan Engstrom, Andrew Ilyas et al.ICML 2020 · 146 citations
- Elo Uncovered: Robustness and Best Practices in Language Model EvaluationMeriem Boubdir, Edward Kim, Beyza Ermis, Sara Hooker et al.NeurIPS 2024 · 94 citations
- InfoLM: A New Metric to Evaluate Summarization & Data2Text GenerationPierre Jean A. Colombo, Chloé Clavel, Pablo PiantanidaAAAI 2022 · 52 citations
Related papers
- What are the best Systems? New Perspectives on NLP BenchmarkingPierre Colombo, Nathan Noiry, Ekhine Irurozki, Stéphan ClémençonNeurIPS 2022 · 20 citations
- Exploiting MMD and Sinkhorn Divergences for Fair and Transferable Representation LearningLuca Oneto, Michele Donini, Giulia Luise, Carlo Ciliberto et al.NeurIPS 2020 · 56 citations
- Selective Preference AggregationShreyas Kadekodi, Hayden McTavish, Berk UstunICML 2025
- Foundations of the Theory of Performance-Based RankingSébastien Piérard, Anaïs Halin, Anthony Cioppa, Adrien Deliège et al.CVPR 2025
- When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model LeaderboardsNorah A. Alzahrani, Hisham Abdullah Alyahya, Yazeed Alnumay, Sultan Alrashed et al.ACL 2024 · 13 citations
