Probabilistic Bilevel Coreset Selection
Xiao Zhou, Renjie Pi, Weizhong Zhang, Yong Lin, Zonghao Chen, Tong Zhang
Abstract
The goal of coreset selection in supervised learning is to produce a weighted subset of data, so that training only on the subset achieves similar performance as training on the entire dataset. Existing methods achieved promising results in resource-constrained scenarios such as continual learning and streaming. However, most of the existing algorithms are limited to traditional machine learning models. A few algorithms that can handle large models adopt greedy search approaches due to the difficulty in solving the discrete subset selection problem, which is computationally costly when coreset becomes larger and often produces suboptimal results. In this work, for the first time we propose a continuous probabilistic bilevel formulation of coreset selection by learning a probablistic weight for each training sample. The overall objective is posed as a bilevel optimization problem, where 1) the inner loop samples coresets and train the model to convergence and 2) the outer loop updates the sample probability progressively according to the model's performance. Importantly, we develop an efficient solver to the bilevel optimization problem via unbiased policy gradient without trouble of implicit differentiation. We provide the convergence property of our training procedure and demonstrate the superiority of our algorithm against various coreset selection methods in various tasks, especially in more challenging label-noise and class-imbalance scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ddfb5760-c3b1-48dd-affe-ceb95e64648eCited by top-tier papers22
- On Penalty-based Bilevel Gradient Descent MethodHan Shen, Tianyi ChenICML 2023 · 105 citations
- DOGE: Domain Reweighting with Generalization EstimationSimin Fan, Matteo Pagliardini, Martin JaggiICML 2024 · 79 citations
- Model Agnostic Sample Reweighting for Out-of-Distribution LearningXiao Zhou, Yong Lin, Renjie Pi, Weizhong Zhang et al.ICML 2022 · 73 citations
- Bilevel Coreset Selection in Continual Learning: A New Formulation and AlgorithmJie Hao, Kaiyi Ji, Mingrui LiuNeurIPS 2023 · 43 citations
- How Benchmark Prediction from Fewer Data Misses the MarkGuanhua Zhang, Florian E. Dorner, Moritz HardtNeurIPS 2025 · 26 citations
Builds on14
- Coresets via Bilevel Optimization for Continual Learning and StreamingZalán Borsos, Mojmir Mutny, Andreas KrauseNeurIPS 2020 · 320 citations
- Dataset Meta-Learning from Kernel Ridge-RegressionTimothy Nguyen, Zhourong Chen, Jaehoon LeeICLR 2021 · 307 citations
- On the Iteration Complexity of Hypergradient ComputationRiccardo Grazzi, Luca Franceschi, Massimiliano Pontil, Saverio SalzoICML 2020 · 241 citations
- Online Continual Learning from Imbalanced DataAristotelis Chrysakis, Marie-Francine MoensICML 2020 · 166 citations
- Bridging the Gap between Sample-based and One-shot Neural Architecture Search with BONASHan Shi, Renjie Pi, Hang Xu, Zhenguo Li et al.NeurIPS 2020 · 148 citations
Related papers
- Coreset Selection via Reducible Loss in Continual LearningRuilin Tong, Yuhang Liu, Javen Qinfeng Shi, Dong GongICLR 2025
- Online Coreset Selection for Rehearsal-based Continual LearningJaehong Yoon, Divyam Madaan, Eunho Yang, Sung Ju HwangICLR 2022 · 181 citations
- GCR: Gradient Coreset based Replay Buffer Selection for Continual LearningRishabh Tiwari, KrishnaTeja Killamsetty, Rishabh K. Iyer, Pradeep ShenoyCVPR 2022 · 102 citations
- Robust and Fully-Dynamic Coreset for Continuous-and-Bounded Learning (With Outliers) ProblemsZixiu Wang, Yiwen Guo, Hu DingNeurIPS 2021 · 10 citations
- RETRIEVE: Coreset Selection for Efficient and Robust Semi-Supervised LearningKrishnaTeja Killamsetty, Xujiang Zhao, Feng Chen, Rishabh K. IyerNeurIPS 2021 · 115 citations
