A Theory of Dynamic Benchmarks
Ali Shirali, Rediet Abebe, Moritz Hardt
Abstract
Dynamic benchmarks interweave model fitting and data collection in an attempt to mitigate the limitations of static benchmarks. In contrast to an extensive theoretical and empirical study of the static setting, the dynamic counterpart lags behind due to limited empirical studies and no apparent theoretical foundation to date. Responding to this deficit, we initiate a theoretical study of dynamic benchmarking. We examine two realizations, one capturing current practice and the other modeling more complex settings. In the first model, where data collection and model fitting alternate sequentially, we prove that model performance improves initially but can stall after only three rounds. Label noise arising from, for instance, annotator disagreement leads to even stronger negative results. Our second model generalizes the first to the case where data collection and model fitting have a hierarchical dependency structure. We show that this design guarantees strictly more progress than the first, albeit at a significant increase in complexity. We support our theoretical analysis by simulating dynamic benchmarks on two popular datasets. These results illuminate the benefits and practical limitations of dynamic benchmarking, providing both a theoretical foundation and a causal explanation for observed bottlenecks in empirical work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 20a03ec0-8a56-4ec3-b515-0d9158481849Cited by top-tier papers9
- Inherent Trade-Offs between Diversity and Stability in Multi-Task BenchmarksGuanhua Zhang, Moritz HardtICML 2024 · 22 citations
- Statistical Multicriteria Benchmarking via the GSD-FrontChristoph Jansen, Georg Schollmeyer, Julian Rodemann, Hannah Blocher et al.NeurIPS 2024 · 15 citations
- Automating Data Annotation under Strategic Human Agents: Risks and Potential SolutionsTian Xie, Xueru ZhangNeurIPS 2024 · 12 citations
- Efficient Lifelong Model Evaluation in an Era of Rapid ProgressAmeya Prabhu, Vishaal Udandarao, Philip Torr, Matthias Bethge et al.NeurIPS 2024 · 11 citations
- VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-CheckingMark Rothermel, Marcus Kornmann, Marcus Rohrbach, Anna RohrbachACL 2026 · 4 citations
Builds on6
- Measuring Robustness to Natural Distribution Shifts in Image ClassificationRohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini et al.NeurIPS 2020 · 731 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- Adversarial Filters of Dataset BiasesRonan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers et al.ICML 2020 · 242 citations
- Dynaboard: An Evaluation-As-A-Service Platform for Holistic Next-Generation BenchmarkingZhiyi Ma, Kawin Ethayarajh, Tristan Thrush, Somya Jain et al.NeurIPS 2021 · 76 citations
- On the Efficacy of Adversarial Data Collection for Question Answering: Results from a Large-Scale Randomized StudyDivyansh Kaushik, Douwe Kiela, Zachary C. Lipton, Wen-tau YihACL 2021
Related papers
- KaLOS finds Consensus: A Meta-Algorithm for Evaluating Inter-Annotator Agreement in Complex Vision TasksDavid Tschirschwitz, Volker RodehorstCVPR 2026 · 1 citation
- Benchmarking Large Language Models Under Data Contamination: A Survey from Static to Dynamic EvaluationSimin Chen, Yiming Chen, Zexin Li, Yifan Jiang et al.EMNLP 2025 · 2 citations
- Leaderboard Incentives: Model Rankings under Strategic Post-TrainingYatong Chen, Guanhua Zhang, Moritz HardtICML 2026 · 2 citations
- Do Question Answering Modeling Improvements Hold Across Benchmarks?Nelson F. Liu, Tony Lee, Robin Jia, Percy LiangACL 2023 · 1 citation
- Noise Correction on Subjective DatasetsUthman Jinadu, Yi DingACL 2024 · 2 citations
