How Much More Data Do I Need? Estimating Requirements for Downstream Tasks
Rafid Mahmood, James Lucas, David Acuna, Daiqing Li, Jonah Philion, José M. Álvarez, Zhiding Yu, Sanja Fidler, Marc T. Law
摘要
Given a small training data set and a learning algorithm, how much more data is necessary to reach a target validation or test performance? This question is of critical importance in applications such as autonomous driving or medical imaging where collecting data is expensive and time-consuming. Overestimating or underestimating data requirements incurs substantial costs that could be avoided with an adequate budget. Prior work on neural scaling laws suggest that the power-law function can fit the validation performance curve and extrapolate it to larger data set sizes. We find that this does not immediately translate to the more difficult downstream task of estimating the required data set size to meet a target performance. In this work, we consider a broad class of computer vision tasks and systematically investigate a family of functions that generalize the power-law function to allow for better estimation of data requirements. Finally, we show that incorporating a tuned correction factor and collecting over multiple rounds significantly improves the performance of the data estimators. Using our guidelines, practitioners can accurately estimate data requirements of machine learning systems to gain savings in both development time and data acquisition costs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Too Large; Data Reduction for Vision-Language Pre-TrainingAlex Jinpeng Wang, Kevin Qinghong Lin, David Junhao Zhang, Stan Weixian Lei 等ICCV 2023 · 被引用 35 次
- Performance Scaling via Optimal Transport: Enabling Data Selection from Partially Revealed SourcesFeiyang Kang, Hoang Anh Just, Anit Kumar Sahu, Ruoxi JiaNeurIPS 2023 · 被引用 21 次
- Delegated ClassificationEden Saig, Inbal Talgam-Cohen, Nir RosenfeldNeurIPS 2023 · 被引用 19 次
- ViCTr: Vital Consistency Transfer for Pathology Aware Image SynthesisOnkar Susladkar, Gayatri Deshmukh, Yalcin Tur, Gorkem Durak 等ICCV 2025 · 被引用 2 次
- Breaking Neural Network Scaling Laws with ModularityAkhilan Boopathy, Sunshine Jiang, William Yue, Jaedong Hwang 等ICLR 2025 · 被引用 1 次
它引用的顶会 Paper11
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 被引用 767 次
- Variational Adversarial Active LearningSamarth Sinha, Sayna Ebrahimi, Trevor DarrellICCV 2019 · 被引用 662 次
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman 等ICLR 2020 · 被引用 462 次
- A Constructive Prediction of the Generalization Error Across ScalesJonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, Nir ShavitICLR 2020 · 被引用 265 次
- Exploring the Limits of Large Scale Pre-trainingSamira Abnar, Mostafa Dehghani, Behnam Neyshabur, Hanie SedghiICLR 2022 · 被引用 135 次
相关 Paper
- A Meta-Learning Approach to Predicting Performance and Data RequirementsAchin Jain, Gurumurthy Swaminathan, Paolo Favaro, Hao Yang 等CVPR 2023
- Optimizing Data Collection for Machine LearningRafid Mahmood, James Lucas, José M. Álvarez, Sanja Fidler 等NeurIPS 2022 · 被引用 39 次
- Model Performance Scaling with Multiple Data SourcesTatsunori HashimotoICML 2021 · 被引用 38 次
- Revisiting Neural Scaling Laws in Language and VisionIbrahim M. Alabdulmohsin, Behnam Neyshabur, Xiaohua ZhaiNeurIPS 2022 · 被引用 171 次
- Scaling Laws for the Value of Individual Data Points in Machine LearningIan Connick Covert, Wenlong Ji, Tatsunori Hashimoto, James ZouICML 2024 · 被引用 12 次
