Adaptive Second Order Coresets for Data-efficient Machine Learning
Omead Pooladzandi, David Davini, Baharan Mirzasoleiman
摘要
Training machine learning models on massive datasets incurs substantial computational costs. To alleviate such costs, there has been a sustained effort to develop data-efficient training methods that can carefully select subsets of the training examples that generalize on par with the full training data. However, existing methods are limited in providing theoretical guarantees for the quality of the models trained on the extracted subsets, and may perform poorly in practice. We propose ADACORE, a method that leverages the geometry of the data to extract subsets of the training examples for efficient machine learning. The key idea behind our method is to dynamically approximate the curvature of the loss function via an exponentially-averaged estimate of the Hessian to select weighted subsets (coresets) that provide a close approximation of the full gradient preconditioned with the Hessian. We prove rigorous guarantees for the convergence of various first and second-order methods applied to the subsets chosen by ADACORE. Our extensive experiments show that ADACORE extracts coresets with higher quality compared to baselines and speeds up training of convex and non-convex machine learning models, such as logistic regression and neural networks, by over 2.9x over the full data and 4.5x over random subsets 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper35
- M3D: Dataset Condensation by Minimizing Maximum Mean DiscrepancyHansong Zhang, Shikun Li, Pengju Wang, Dan Zeng 等AAAI 2024 · 被引用 63 次
- SmallToLarge (S2L): Scalable Data Selection for Fine-tuning Large Language Models by Summarizing Training Trajectories of Small ModelsYu Yang, Siddhartha Mishra, Jeffrey N. Chiang, Baharan MirzasoleimanNeurIPS 2024 · 被引用 63 次
- Towards Sustainable Learning: Coresets for Data-efficient Deep LearningYu Yang, Hao Kang, Baharan MirzasoleimanICML 2023 · 被引用 58 次
- Dataset Distillation with Convexified Implicit GradientsNoel Loo, Ramin M. Hasani, Mathias Lechner, Daniela RusICML 2023 · 被引用 56 次
- You Only Condense Once: Two Rules for Pruning Condensed DatasetsYang He, Lingao Xiao, Joey Tianyi ZhouNeurIPS 2023 · 被引用 33 次
它引用的顶会 Paper3
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 被引用 494 次
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman 等ICLR 2020 · 被引用 462 次
- BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask LearningFisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian 等CVPR 2020
相关 Paper
- Efficient Coreset Selection with Cluster-based MethodsChengliang Chai, Jiayi Wang, Nan Tang, Ye Yuan 等KDD 2023 · 被引用 19 次
- GRAD-MATCH: Gradient Matching based Data Subset Selection for Efficient Deep Model TrainingKrishnaTeja Killamsetty, Durga Sivasubramanian, Ganesh Ramakrishnan, Abir De 等ICML 2021 · 被引用 305 次
- Samples with Low Loss Curvature Improve Data EfficiencyIsha Garg, Kaushik RoyCVPR 2023
- Coresets over Multiple Tables for Feature-rich and Data-efficient Machine LearningJiayi Wang, Chengliang Chai, Nan Tang, Jiabin Liu 等VLDB 2023 · 被引用 31 次
- Mind the Boundary: Coreset Selection via Reconstructing the Decision BoundaryShuo Yang, Zhe Cao, Sheng Guo, Ruiheng Zhang 等ICML 2024 · 被引用 25 次
