Optimizing Data Collection for Machine Learning
Rafid Mahmood, James Lucas, José M. Álvarez, Sanja Fidler, Marc T. Law
摘要
Modern deep learning systems require huge data sets to achieve impressive performance, but there is little guidance on how much or what kind of data to collect. Over-collecting data incurs unnecessary present costs, while under-collecting may incur future costs and delay workflows. We propose a new paradigm for modeling the data collection workflow as a formal optimal data collection problem that allows designers to specify performance targets, collection costs, a time horizon, and penalties for failing to meet the targets. Additionally, this formulation generalizes to tasks requiring multiple data sources, such as labeled and unlabeled data used in semi-supervised learning. To solve our problem, we develop Learn-Optimize-Collect (LOC), which minimizes expected future collection costs. Finally, we numerically compare our framework to the conventional baseline of estimating data requirements by extrapolating from neural scaling laws. We significantly reduce the risks of failing to meet desired performance targets on several classification, segmentation, and detection tasks, while maintaining low total collection costs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Scaling Laws for Hyperparameter OptimizationArlind Kadra, Maciej Janowski, Martin Wistuba, Josif GrabockaNeurIPS 2023 · 被引用 23 次
- How to Select Which Active Learning Strategy is Best Suited for Your Specific Problem and BudgetGuy Hacohen, Daphna WeinshallNeurIPS 2023 · 被引用 23 次
- Pricing and Competition for Generative AIRafid MahmoodNeurIPS 2024 · 被引用 18 次
- Scaling-Aware Data Selection for End-to-End Autonomous Driving SystemsTolga Dimlioglu, Nadine Chang, Maying Shen, Rafid Mahmood 等CVPR 2026 · 被引用 1 次
- Adaptive Data Collection for Robust Learning Across Multiple DistributionsChengbo Zang, Mehmet Kerem Türkcan, Gil Zussman, Zoran Kostic 等ICML 2025
它引用的顶会 Paper11
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 被引用 767 次
- Variational Adversarial Active LearningSamarth Sinha, Sayna Ebrahimi, Trevor DarrellICCV 2019 · 被引用 662 次
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman 等ICLR 2020 · 被引用 462 次
- A Constructive Prediction of the Generalization Error Across ScalesJonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, Nir ShavitICLR 2020 · 被引用 265 次
- Not All Labels Are Equal: Rationalizing The Labeling Costs for Training Object DetectionIsmail Elezi, Zhiding Yu, Anima Anandkumar, Laura Leal-Taixé 等CVPR 2022 · 被引用 45 次
相关 Paper
- How Much More Data Do I Need? Estimating Requirements for Downstream TasksRafid Mahmood, James Lucas, David Acuna, Daiqing Li 等CVPR 2022 · 被引用 21 次
- Navigating Scaling Laws: Compute Optimality in Adaptive Model TrainingSotiris Anagnostidis, Gregor Bachmann, Imanol Schlag, Thomas HofmannICML 2024 · 被引用 2 次
- Model Performance Scaling with Multiple Data SourcesTatsunori HashimotoICML 2021 · 被引用 38 次
- Leverage Class-Specific Accuracy to Guide Data Generation for Improving Image ClassificationJay Gala, Pengtao XieICML 2024 · 被引用 1 次
- UniT: Unified Knowledge Transfer for Any-Shot Object Detection and SegmentationSiddhesh Khandelwal, Raghav Goyal, Leonid SigalCVPR 2021
