Inherent Trade-Offs between Diversity and Stability in Multi-Task Benchmarks
Guanhua Zhang, Moritz Hardt
摘要
We examine multi-task benchmarks in machine learning through the lens of social choice theory. We draw an analogy between benchmarks and electoral systems, where models are candidates and tasks are voters. This suggests a distinction between cardinal and ordinal benchmark systems. The former aggregate numerical scores into one model ranking; the latter aggregate rankings for each task. We apply Arrow's impossibility theorem to ordinal benchmarks to highlight the inherent limitations of ordinal systems, particularly their sensitivity to the inclusion of irrelevant models. Inspired by Arrow's theorem, we empirically demonstrate a strong trade-off between diversity and sensitivity to irrelevant changes in existing multi-task benchmarks. Our result is based on new quantitative measures of diversity and sensitivity that we introduce. Sensitivity quantifies the impact that irrelevant changes to tasks have on a benchmark. Diversity captures the degree of disagreement in model rankings across tasks. We develop efficient approximation algorithms for both measures, as exact computation is computationally challenging. Through extensive experiments on seven cardinal benchmarks and eleven ordinal benchmarks, we demonstrate a clear trade-off between diversity and stability: The more diverse a multi-task benchmark, the more sensitive to trivial changes it is. Additionally, we show that the aggregated rankings of existing benchmarks are highly unstable under irrelevant changes. The codes and data are available at https://socialfoundations. github.io/benchbench/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Train-before-Test Harmonizes Language Model RankingsGuanhua Zhang, Ricardo Dominguez-Olmedo, Moritz HardtICLR 2026 · 被引用 13 次
- Multivariate Stochastic Dominance via Optimal Transport and Applications to Models BenchmarkingGabriel Rioux, Apoorva Nitsure, Mattia Rigotti, Kristjan H. Greenewald 等NeurIPS 2024 · 被引用 6 次
- Nonparametric LLM Evaluation from Preference DataDennis Frauen, Athiya Deviyani, Mihaela van der Schaar, Stefan FeuerriegelICML 2026 · 被引用 5 次
- Metritocracy: Representative Metrics for Lite BenchmarksAriel D. Procaccia, Ben Schiffer, Serena Wang, Shirley ZhangNeurIPS 2025 · 被引用 2 次
- Robust AI Evaluation through Maximal LotteriesHadi Khalaf, Serena Wang, Daniel Halpern, Itai Shapira 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper10
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- From ImageNet to Image Classification: Contextualizing Progress on BenchmarksDimitris Tsipras, Shibani Santurkar, Logan Engstrom, Andrew Ilyas 等ICML 2020 · 被引用 146 次
- Elo Uncovered: Robustness and Best Practices in Language Model EvaluationMeriem Boubdir, Edward Kim, Beyza Ermis, Sara Hooker 等NeurIPS 2024 · 被引用 94 次
- InfoLM: A New Metric to Evaluate Summarization & Data2Text GenerationPierre Jean A. Colombo, Chloé Clavel, Pablo PiantanidaAAAI 2022 · 被引用 52 次
相关 Paper
- What are the best Systems? New Perspectives on NLP BenchmarkingPierre Colombo, Nathan Noiry, Ekhine Irurozki, Stéphan ClémençonNeurIPS 2022 · 被引用 20 次
- Exploiting MMD and Sinkhorn Divergences for Fair and Transferable Representation LearningLuca Oneto, Michele Donini, Giulia Luise, Carlo Ciliberto 等NeurIPS 2020 · 被引用 56 次
- Selective Preference AggregationShreyas Kadekodi, Hayden McTavish, Berk UstunICML 2025
- Foundations of the Theory of Performance-Based RankingSébastien Piérard, Anaïs Halin, Anthony Cioppa, Adrien Deliège 等CVPR 2025
- When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model LeaderboardsNorah A. Alzahrani, Hisham Abdullah Alyahya, Yazeed Alnumay, Sultan Alrashed 等ACL 2024 · 被引用 13 次
