How Much More Data Do I Need? Estimating Requirements for Downstream Tasks
Rafid Mahmood, James Lucas, David Acuna, Daiqing Li, Jonah Philion, José M. Álvarez, Zhiding Yu, Sanja Fidler, Marc T. Law
Abstract
Given a small training data set and a learning algorithm, how much more data is necessary to reach a target validation or test performance? This question is of critical importance in applications such as autonomous driving or medical imaging where collecting data is expensive and time-consuming. Overestimating or underestimating data requirements incurs substantial costs that could be avoided with an adequate budget. Prior work on neural scaling laws suggest that the power-law function can fit the validation performance curve and extrapolate it to larger data set sizes. We find that this does not immediately translate to the more difficult downstream task of estimating the required data set size to meet a target performance. In this work, we consider a broad class of computer vision tasks and systematically investigate a family of functions that generalize the power-law function to allow for better estimation of data requirements. Finally, we show that incorporating a tuned correction factor and collecting over multiple rounds significantly improves the performance of the data estimators. Using our guidelines, practitioners can accurately estimate data requirements of machine learning systems to gain savings in both development time and data acquisition costs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 085ac1ff-fd98-4cc0-ae79-b696377fd725Cited by top-tier papers9
- Too Large; Data Reduction for Vision-Language Pre-TrainingAlex Jinpeng Wang, Kevin Qinghong Lin, David Junhao Zhang, Stan Weixian Lei et al.ICCV 2023 · 35 citations
- Performance Scaling via Optimal Transport: Enabling Data Selection from Partially Revealed SourcesFeiyang Kang, Hoang Anh Just, Anit Kumar Sahu, Ruoxi JiaNeurIPS 2023 · 21 citations
- Delegated ClassificationEden Saig, Inbal Talgam-Cohen, Nir RosenfeldNeurIPS 2023 · 19 citations
- ViCTr: Vital Consistency Transfer for Pathology Aware Image SynthesisOnkar Susladkar, Gayatri Deshmukh, Yalcin Tur, Gorkem Durak et al.ICCV 2025 · 2 citations
- Breaking Neural Network Scaling Laws with ModularityAkhilan Boopathy, Sunshine Jiang, William Yue, Jaedong Hwang et al.ICLR 2025 · 1 citation
Builds on11
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
- Variational Adversarial Active LearningSamarth Sinha, Sayna Ebrahimi, Trevor DarrellICCV 2019 · 662 citations
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman et al.ICLR 2020 · 462 citations
- A Constructive Prediction of the Generalization Error Across ScalesJonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, Nir ShavitICLR 2020 · 265 citations
- Exploring the Limits of Large Scale Pre-trainingSamira Abnar, Mostafa Dehghani, Behnam Neyshabur, Hanie SedghiICLR 2022 · 135 citations
Related papers
- A Meta-Learning Approach to Predicting Performance and Data RequirementsAchin Jain, Gurumurthy Swaminathan, Paolo Favaro, Hao Yang et al.CVPR 2023
- Optimizing Data Collection for Machine LearningRafid Mahmood, James Lucas, José M. Álvarez, Sanja Fidler et al.NeurIPS 2022 · 39 citations
- Model Performance Scaling with Multiple Data SourcesTatsunori HashimotoICML 2021 · 38 citations
- Revisiting Neural Scaling Laws in Language and VisionIbrahim M. Alabdulmohsin, Behnam Neyshabur, Xiaohua ZhaiNeurIPS 2022 · 171 citations
- Scaling Laws for the Value of Individual Data Points in Machine LearningIan Connick Covert, Wenlong Ji, Tatsunori Hashimoto, James ZouICML 2024 · 12 citations
