Optimizing Data Collection for Machine Learning
Rafid Mahmood, James Lucas, José M. Álvarez, Sanja Fidler, Marc T. Law
Abstract
Modern deep learning systems require huge data sets to achieve impressive performance, but there is little guidance on how much or what kind of data to collect. Over-collecting data incurs unnecessary present costs, while under-collecting may incur future costs and delay workflows. We propose a new paradigm for modeling the data collection workflow as a formal optimal data collection problem that allows designers to specify performance targets, collection costs, a time horizon, and penalties for failing to meet the targets. Additionally, this formulation generalizes to tasks requiring multiple data sources, such as labeled and unlabeled data used in semi-supervised learning. To solve our problem, we develop Learn-Optimize-Collect (LOC), which minimizes expected future collection costs. Finally, we numerically compare our framework to the conventional baseline of estimating data requirements by extrapolating from neural scaling laws. We significantly reduce the risks of failing to meet desired performance targets on several classification, segmentation, and detection tasks, while maintaining low total collection costs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Scaling Laws for Hyperparameter OptimizationArlind Kadra, Maciej Janowski, Martin Wistuba, Josif GrabockaNeurIPS 2023 · 23 citations
- How to Select Which Active Learning Strategy is Best Suited for Your Specific Problem and BudgetGuy Hacohen, Daphna WeinshallNeurIPS 2023 · 23 citations
- Pricing and Competition for Generative AIRafid MahmoodNeurIPS 2024 · 18 citations
- Scaling-Aware Data Selection for End-to-End Autonomous Driving SystemsTolga Dimlioglu, Nadine Chang, Maying Shen, Rafid Mahmood et al.CVPR 2026 · 1 citation
- Adaptive Data Collection for Robust Learning Across Multiple DistributionsChengbo Zang, Mehmet Kerem Türkcan, Gil Zussman, Zoran Kostic et al.ICML 2025
Builds on11
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
- Variational Adversarial Active LearningSamarth Sinha, Sayna Ebrahimi, Trevor DarrellICCV 2019 · 662 citations
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman et al.ICLR 2020 · 462 citations
- A Constructive Prediction of the Generalization Error Across ScalesJonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, Nir ShavitICLR 2020 · 265 citations
- Not All Labels Are Equal: Rationalizing The Labeling Costs for Training Object DetectionIsmail Elezi, Zhiding Yu, Anima Anandkumar, Laura Leal-Taixé et al.CVPR 2022 · 45 citations
Related papers
- How Much More Data Do I Need? Estimating Requirements for Downstream TasksRafid Mahmood, James Lucas, David Acuna, Daiqing Li et al.CVPR 2022 · 21 citations
- Navigating Scaling Laws: Compute Optimality in Adaptive Model TrainingSotiris Anagnostidis, Gregor Bachmann, Imanol Schlag, Thomas HofmannICML 2024 · 2 citations
- Model Performance Scaling with Multiple Data SourcesTatsunori HashimotoICML 2021 · 38 citations
- Leverage Class-Specific Accuracy to Guide Data Generation for Improving Image ClassificationJay Gala, Pengtao XieICML 2024 · 1 citation
- UniT: Unified Knowledge Transfer for Any-Shot Object Detection and SegmentationSiddhesh Khandelwal, Raghav Goyal, Leonid SigalCVPR 2021
