A Theory of Dynamic Benchmarks
Ali Shirali, Rediet Abebe, Moritz Hardt
摘要
Dynamic benchmarks interweave model fitting and data collection in an attempt to mitigate the limitations of static benchmarks. In contrast to an extensive theoretical and empirical study of the static setting, the dynamic counterpart lags behind due to limited empirical studies and no apparent theoretical foundation to date. Responding to this deficit, we initiate a theoretical study of dynamic benchmarking. We examine two realizations, one capturing current practice and the other modeling more complex settings. In the first model, where data collection and model fitting alternate sequentially, we prove that model performance improves initially but can stall after only three rounds. Label noise arising from, for instance, annotator disagreement leads to even stronger negative results. Our second model generalizes the first to the case where data collection and model fitting have a hierarchical dependency structure. We show that this design guarantees strictly more progress than the first, albeit at a significant increase in complexity. We support our theoretical analysis by simulating dynamic benchmarks on two popular datasets. These results illuminate the benefits and practical limitations of dynamic benchmarking, providing both a theoretical foundation and a causal explanation for observed bottlenecks in empirical work.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Inherent Trade-Offs between Diversity and Stability in Multi-Task BenchmarksGuanhua Zhang, Moritz HardtICML 2024 · 被引用 22 次
- Statistical Multicriteria Benchmarking via the GSD-FrontChristoph Jansen, Georg Schollmeyer, Julian Rodemann, Hannah Blocher 等NeurIPS 2024 · 被引用 15 次
- Automating Data Annotation under Strategic Human Agents: Risks and Potential SolutionsTian Xie, Xueru ZhangNeurIPS 2024 · 被引用 12 次
- Efficient Lifelong Model Evaluation in an Era of Rapid ProgressAmeya Prabhu, Vishaal Udandarao, Philip Torr, Matthias Bethge 等NeurIPS 2024 · 被引用 11 次
- VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-CheckingMark Rothermel, Marcus Kornmann, Marcus Rohrbach, Anna RohrbachACL 2026 · 被引用 4 次
它引用的顶会 Paper6
- Measuring Robustness to Natural Distribution Shifts in Image ClassificationRohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini 等NeurIPS 2020 · 被引用 731 次
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal 等ACL 2020 · 被引用 602 次
- Adversarial Filters of Dataset BiasesRonan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers 等ICML 2020 · 被引用 242 次
- Dynaboard: An Evaluation-As-A-Service Platform for Holistic Next-Generation BenchmarkingZhiyi Ma, Kawin Ethayarajh, Tristan Thrush, Somya Jain 等NeurIPS 2021 · 被引用 76 次
- On the Efficacy of Adversarial Data Collection for Question Answering: Results from a Large-Scale Randomized StudyDivyansh Kaushik, Douwe Kiela, Zachary C. Lipton, Wen-tau YihACL 2021
相关 Paper
- KaLOS finds Consensus: A Meta-Algorithm for Evaluating Inter-Annotator Agreement in Complex Vision TasksDavid Tschirschwitz, Volker RodehorstCVPR 2026 · 被引用 1 次
- Benchmarking Large Language Models Under Data Contamination: A Survey from Static to Dynamic EvaluationSimin Chen, Yiming Chen, Zexin Li, Yifan Jiang 等EMNLP 2025 · 被引用 2 次
- Leaderboard Incentives: Model Rankings under Strategic Post-TrainingYatong Chen, Guanhua Zhang, Moritz HardtICML 2026 · 被引用 2 次
- Do Question Answering Modeling Improvements Hold Across Benchmarks?Nelson F. Liu, Tony Lee, Robin Jia, Percy LiangACL 2023 · 被引用 1 次
- Noise Correction on Subjective DatasetsUthman Jinadu, Yi DingACL 2024 · 被引用 2 次
