Feature Distribution Matching by Optimal Transport for Effective and Robust Coreset Selection
Weiwei Xiao, Yongyong Chen, Qiben Shan, Yaowei Wang, Jingyong Su
Abstract
Training neural networks with good generalization requires large computational costs in many deep learning methods due to large-scale datasets and over-parameterized models. Despite the emergence of a number of coreset selection methods to reduce the computational costs, the problem of coreset distribution bias, i.e., the skewed distribution between the coreset and the entire dataset, has not been well studied. In this paper, we find that the closer the feature distribution of the coreset is to that of the entire dataset, the better the generalization performance of the coreset, particularly under extreme pruning. This motivates us to propose a simple yet effective method for coreset selection to alleviate the distribution bias between the coreset and the entire dataset, called feature distribution matching (FDMat). Unlike gradient-based methods, which selects samples with larger gradient values or approximates gradient values of the entire dataset, FD-Mat aims to select the coreset that is closest to the feature distribution of the entire dataset. Specifically, FDMat transfers coreset selection as an optimal transport problem from the coreset to the entire dataset in feature embedding spaces. Moreover, our method shows strong robustness due to the removal of samples far from the distribution, especially for the entire dataset containing noisy and class-imbalanced samples. Extensive experiments on multiple benchmarks show that FDMat can improve the performance of coreset selection than existing coreset methods. The code is available at https://github.com/successhaha/FDMat.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9952d353-232d-44ae-bf69-9afdb8f053aaCited by top-tier papers7
- Selection of LLM Fine-Tuning Data Based on Orthogonal RulesXiaomin Li, Mingye Gao, Zhiwei Zhang, Chang Yue et al.AAAI 2026 · 9 citations
- Datamap-Driven Tabular Coreset Selection for Classifier TrainingAviv Hadar, Tova Milo, Kathy RazmadzeVLDB 2025 · 6 citations
- GORACS: Group-level Optimal Transport-guided Coreset Selection for LLM-based Recommender SystemsTiehua Mei, Hengrui Chen, Peng Yu, Jiaqing Liang et al.KDD 2025 · 3 citations
- TAROT: Targeted Data Selection via Optimal TransportLan Feng, Fan Nie, Yuejiang Liu, Alexandre AlahiICML 2025
- PEARL: Towards Permutation-Resilient LLMsLiang Chen, Li Shen, Yang Deng, Xiaoyan Zhao et al.ICLR 2025
Builds on11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
Related papers
- FedCS: Coreset Selection for Federated LearningChenhe Hao, Weiying Xie, Daixun Li, Haonan Qin et al.CVPR 2025
- Improved Distribution Matching for Dataset CondensationGanlong Zhao, Guanbin Li, Yipeng Qin, Yizhou YuCVPR 2023
- Mind the Boundary: Coreset Selection via Reconstructing the Decision BoundaryShuo Yang, Zhe Cao, Sheng Guo, Ruiheng Zhang et al.ICML 2024 · 25 citations
- FAST: Topology-Aware Frequency-Domain Distribution Matching for Coreset SelectionJin Cui, Boran Zhao, Jiajun Xu, Jiaqi Guo et al.CVPR 2026 · 2 citations
- Learning to Re-weight Examples with Optimal Transport for Imbalanced ClassificationDandan Guo, Zhuo Li, Meixi Zheng, He Zhao et al.NeurIPS 2022 · 46 citations
