Towards Sustainable Learning: Coresets for Data-efficient Deep Learning
Yu Yang, Hao Kang, Baharan Mirzasoleiman
摘要
To improve the efficiency and sustainability of learning deep models, we propose CREST, the first scalable framework with rigorous theoretical guarantees to identify the most valuable examples for training non-convex models, particularly deep networks. To guarantee convergence to a stationary point of a non-convex function, CREST models the non-convex loss as a series of quadratic functions and extracts a coreset for each quadratic sub-region. In addition, to ensure faster convergence of stochastic gradient methods such as (mini-batch) SGD, CREST iteratively extracts multiple mini-batch coresets from larger random subsets of training data, to ensure nearly-unbiased gradients with small variances. Finally, to further improve scalability and efficiency, CREST identifies and excludes the examples that are learned from the coreset selection pipeline. Our extensive experiments on several deep networks trained on vision and NLP datasets, including CIFAR-10, CIFAR-100, TinyImageNet, and SNLI, confirm that CREST speeds up training deep networks on very large datasets, by 1.7x to 2.5x with minimum loss in the performance. By analyzing the learning difficulty of the subsets selected by CREST, we show that deep models benefit the most by learning from subsets of increasing difficulty levels 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- SmallToLarge (S2L): Scalable Data Selection for Fine-tuning Large Language Models by Summarizing Training Trajectories of Small ModelsYu Yang, Siddhartha Mishra, Jeffrey N. Chiang, Baharan MirzasoleimanNeurIPS 2024 · 被引用 63 次
- Data Distillation Can Be Like Vodka: Distilling More Times For Better QualityXuxi Chen, Yu Yang, Zhangyang Wang, Baharan MirzasoleimanICLR 2024 · 被引用 19 次
- D2 Pruning: Message Passing for Balancing Diversity & Difficulty in Data PruningAdyasha Maharana, Prateek Yadav, Mohit BansalICLR 2024 · 被引用 19 次
- Selectivity Drives Productivity: Efficient Dataset Pruning for Enhanced Transfer LearningYihua Zhang, Yimeng Zhang, Aochuan Chen, Jinghan Jia 等NeurIPS 2023 · 被引用 18 次
- Optimizing Data Acquisition to Enhance Machine Learning PerformanceTingting Wang, Shixun Huang, Zhifeng Bao, J. Shane Culpepper 等VLDB 2024 · 被引用 13 次
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 被引用 806 次
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 被引用 767 次
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 被引用 494 次
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman 等ICLR 2020 · 被引用 462 次
相关 Paper
- Adaptive Second Order Coresets for Data-efficient Machine LearningOmead Pooladzandi, David Davini, Baharan MirzasoleimanICML 2022 · 被引用 83 次
- Refined Coreset Selection: Towards Minimal Coreset Size under Model Performance ConstraintsXiaobo Xia, Jiale Liu, Shaokun Zhang, Qingyun Wu 等ICML 2024 · 被引用 17 次
- Samples with Low Loss Curvature Improve Data EfficiencyIsha Garg, Kaushik RoyCVPR 2023
- Stable Coresets via Posterior Sampling: Aligning Induced and Full Loss LandscapesWei-Kai Chang, Rajiv KhannaNeurIPS 2025
- Efficient Representativeness-Aware Coreset SelectionZihao Cheng, Binrui Wu, Zhiwei Li, Yuesen Liao 等NeurIPS 2025 · 被引用 1 次
