Contributing Dimension Structure of Deep Feature for Coreset Selection
Zhijing Wan, Zhixiang Wang, Yuran Wang, Zheng Wang, Hongyuan Zhu, Shin'ichi Satoh
Abstract
Coreset selection seeks to choose a subset of crucial training samples for efficient learning. It has gained traction in deep learning, particularly with the surge in training dataset sizes. Sample selection hinges on two main aspects: a sample's representation in enhancing performance and the role of sample diversity in averting overfitting. Existing methods typically measure both the representation and diversity of data based on similarity metrics, such as L2-norm. They have capably tackled representation via distribution matching guided by the similarities of features, gradients, or other information between data. However, the results of effectively diverse sample selection are mired in sub-optimality. This is because the similarity metrics usually simply aggregate dimension similarities without acknowledging disparities among the dimensions that significantly contribute to the final similarity. As a result, they fall short of adequately capturing diversity. To address this, we propose a feature-based diversity constraint, compelling the chosen subset to exhibit maximum diversity. Our key lies in the introduction of a novel Contributing Dimension Structure (CDS) metric. Different from similarity metrics that measure the overall similarity of high-dimensional features, our CDS metric considers not only the reduction of redundancy in feature dimensions, but also the difference between dimensions that contribute significantly to the final similarity. We reveal that existing methods tend to favor samples with similar CDS, leading to a reduced variety of CDS types within the coreset and subsequently hindering model performance. In response, we enhance the performance of five classical selection methods by integrating the CDS constraint. Our experiments on three datasets demonstrate the general effectiveness of the proposed method in boosting existing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cdbbd224-1202-4ba8-b451-0593c9418b27Cited by top-tier papers5
- Adaptive Dataset QuantizationMuquan Li, Dongyang Zhang, Qiang Dong, Xiurui Xie et al.AAAI 2025 · 9 citations
- Balancing Privacy and Performance: A Many-in-One Approach for Image AnonymizationXuemei Jia, Jiawei Du, Hui Wei, Ruinian Xue et al.AAAI 2025 · 2 citations
- UNSEEN: Enhancing Dataset Pruning from a Generalization PerspectiveFurui Xu, Shaobo Wang, Jiajun Zhang, Chenghao Sun et al.AAAI 2026
- Coreset Selection via Reducible Loss in Continual LearningRuilin Tong, Yuhang Liu, Javen Qinfeng Shi, Dong GongICLR 2025
- Foundation Model Insights and a Multi-Model Approach for Superior Fine-Grained One-shot Subset SelectionZhijing Wan, Zhixiang Wang, Zheng Wang, Xin Xu et al.ICML 2025
Builds on13
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman et al.ICLR 2020 · 462 citations
- Online Coreset Selection for Rehearsal-based Continual LearningJaehong Yoon, Divyam Madaan, Eunho Yang, Sung Ju HwangICLR 2022 · 181 citations
- Adaptive Second Order Coresets for Data-efficient Machine LearningOmead Pooladzandi, David Davini, Baharan MirzasoleimanICML 2022 · 83 citations
Related papers
- CSOR: Coreset Selection for Object Re-identification via Class PruningMinyoung Oh, Jae-Young SimICML 2026
- Efficient Core-set Selection for Deep Learning Through Squared Loss MinimizationJianting ChenICML 2025
- HCDS: Hierarchical Clustering for Cold-Start Few-Shot Data SelectionYuhua Zhao, Zhixin Han, Xunzhi Wang, Bitong Luo et al.SIGIR 2025
- FedCS: Coreset Selection for Federated LearningChenhe Hao, Weiying Xie, Daixun Li, Haonan Qin et al.CVPR 2025
- FAST: Topology-Aware Frequency-Domain Distribution Matching for Coreset SelectionJin Cui, Boran Zhao, Jiajun Xu, Jiaqi Guo et al.CVPR 2026 · 2 citations
