Towards Sustainable Learning: Coresets for Data-efficient Deep Learning
Yu Yang, Hao Kang, Baharan Mirzasoleiman
Abstract
To improve the efficiency and sustainability of learning deep models, we propose CREST, the first scalable framework with rigorous theoretical guarantees to identify the most valuable examples for training non-convex models, particularly deep networks. To guarantee convergence to a stationary point of a non-convex function, CREST models the non-convex loss as a series of quadratic functions and extracts a coreset for each quadratic sub-region. In addition, to ensure faster convergence of stochastic gradient methods such as (mini-batch) SGD, CREST iteratively extracts multiple mini-batch coresets from larger random subsets of training data, to ensure nearly-unbiased gradients with small variances. Finally, to further improve scalability and efficiency, CREST identifies and excludes the examples that are learned from the coreset selection pipeline. Our extensive experiments on several deep networks trained on vision and NLP datasets, including CIFAR-10, CIFAR-100, TinyImageNet, and SNLI, confirm that CREST speeds up training deep networks on very large datasets, by 1.7x to 2.5x with minimum loss in the performance. By analyzing the learning difficulty of the subsets selected by CREST, we show that deep models benefit the most by learning from subsets of increasing difficulty levels 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e32f2563-85fe-4b3d-90e7-6f5d4039b75bCited by top-tier papers23
- SmallToLarge (S2L): Scalable Data Selection for Fine-tuning Large Language Models by Summarizing Training Trajectories of Small ModelsYu Yang, Siddhartha Mishra, Jeffrey N. Chiang, Baharan MirzasoleimanNeurIPS 2024 · 63 citations
- Data Distillation Can Be Like Vodka: Distilling More Times For Better QualityXuxi Chen, Yu Yang, Zhangyang Wang, Baharan MirzasoleimanICLR 2024 · 19 citations
- D2 Pruning: Message Passing for Balancing Diversity & Difficulty in Data PruningAdyasha Maharana, Prateek Yadav, Mohit BansalICLR 2024 · 19 citations
- Selectivity Drives Productivity: Efficient Dataset Pruning for Enhanced Transfer LearningYihua Zhang, Yimeng Zhang, Aochuan Chen, Jinghan Jia et al.NeurIPS 2023 · 18 citations
- Optimizing Data Acquisition to Enhance Machine Learning PerformanceTingting Wang, Shixun Huang, Zhifeng Bao, J. Shane Culpepper et al.VLDB 2024 · 13 citations
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman et al.ICLR 2020 · 462 citations
Related papers
- Adaptive Second Order Coresets for Data-efficient Machine LearningOmead Pooladzandi, David Davini, Baharan MirzasoleimanICML 2022 · 83 citations
- Refined Coreset Selection: Towards Minimal Coreset Size under Model Performance ConstraintsXiaobo Xia, Jiale Liu, Shaokun Zhang, Qingyun Wu et al.ICML 2024 · 17 citations
- Samples with Low Loss Curvature Improve Data EfficiencyIsha Garg, Kaushik RoyCVPR 2023
- Stable Coresets via Posterior Sampling: Aligning Induced and Full Loss LandscapesWei-Kai Chang, Rajiv KhannaNeurIPS 2025
- Efficient Representativeness-Aware Coreset SelectionZihao Cheng, Binrui Wu, Zhiwei Li, Yuesen Liao et al.NeurIPS 2025 · 1 citation
