Feature Distribution Matching by Optimal Transport for Effective and Robust Coreset Selection
Weiwei Xiao, Yongyong Chen, Qiben Shan, Yaowei Wang, Jingyong Su
摘要
Training neural networks with good generalization requires large computational costs in many deep learning methods due to large-scale datasets and over-parameterized models. Despite the emergence of a number of coreset selection methods to reduce the computational costs, the problem of coreset distribution bias, i.e., the skewed distribution between the coreset and the entire dataset, has not been well studied. In this paper, we find that the closer the feature distribution of the coreset is to that of the entire dataset, the better the generalization performance of the coreset, particularly under extreme pruning. This motivates us to propose a simple yet effective method for coreset selection to alleviate the distribution bias between the coreset and the entire dataset, called feature distribution matching (FDMat). Unlike gradient-based methods, which selects samples with larger gradient values or approximates gradient values of the entire dataset, FD-Mat aims to select the coreset that is closest to the feature distribution of the entire dataset. Specifically, FDMat transfers coreset selection as an optimal transport problem from the coreset to the entire dataset in feature embedding spaces. Moreover, our method shows strong robustness due to the removal of samples far from the distribution, especially for the entire dataset containing noisy and class-imbalanced samples. Extensive experiments on multiple benchmarks show that FDMat can improve the performance of coreset selection than existing coreset methods. The code is available at https://github.com/successhaha/FDMat.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Selection of LLM Fine-Tuning Data Based on Orthogonal RulesXiaomin Li, Mingye Gao, Zhiwei Zhang, Chang Yue 等AAAI 2026 · 被引用 9 次
- Datamap-Driven Tabular Coreset Selection for Classifier TrainingAviv Hadar, Tova Milo, Kathy RazmadzeVLDB 2025 · 被引用 6 次
- GORACS: Group-level Optimal Transport-guided Coreset Selection for LLM-based Recommender SystemsTiehua Mei, Hengrui Chen, Peng Yu, Jiaqing Liang 等KDD 2025 · 被引用 3 次
- TAROT: Targeted Data Selection via Optimal TransportLan Feng, Fan Nie, Yuejiang Liu, Alexandre AlahiICML 2025
- PEARL: Towards Permutation-Resilient LLMsLiang Chen, Li Shen, Yang Deng, Xiaoyan Zhao 等ICLR 2025
它引用的顶会 Paper11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 被引用 806 次
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 被引用 684 次
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 被引用 494 次
相关 Paper
- FedCS: Coreset Selection for Federated LearningChenhe Hao, Weiying Xie, Daixun Li, Haonan Qin 等CVPR 2025
- Improved Distribution Matching for Dataset CondensationGanlong Zhao, Guanbin Li, Yipeng Qin, Yizhou YuCVPR 2023
- Mind the Boundary: Coreset Selection via Reconstructing the Decision BoundaryShuo Yang, Zhe Cao, Sheng Guo, Ruiheng Zhang 等ICML 2024 · 被引用 25 次
- FAST: Topology-Aware Frequency-Domain Distribution Matching for Coreset SelectionJin Cui, Boran Zhao, Jiajun Xu, Jiaqi Guo 等CVPR 2026 · 被引用 2 次
- Learning to Re-weight Examples with Optimal Transport for Imbalanced ClassificationDandan Guo, Zhuo Li, Meixi Zheng, He Zhao 等NeurIPS 2022 · 被引用 46 次
