Balancing Feature Similarity and Label Variability for Optimal Size-Aware One-shot Subset Selection
Abhinab Acharya, Dayou Yu, Qi Yu, Xumin Liu
Abstract
Subset or core-set selection offers a data-efficient way for training deep learning models. One-shot subset selection poses additional challenges as subset selection is only performed once and full set data become unavailable after the selection. However, most existing methods tend to choose either diverse or difficult data samples, which fail to faithfully represent the joint data distribution that is comprised of both feature and label information. The selection is also performed independently from the subset size, which plays an essential role in choosing what types of samples. To address this critical gap, we propose to conduct Feature similarity and Label variability Balanced One-shot Subset Selection (BOSS), aiming to construct an optimal size-aware subset for data-efficient deep learning. We show that a novel balanced core-set loss bound theoretically justifies the need to simultaneously consider both diversity and difficulty to form an optimal subset. It also reveals how the subset size influences the bound. We further connect the inaccessible bound to a practical surrogate target which is tailored to subset sizes and varying levels of overall difficulty. We design a novel Beta-scoring importance function to delicately control the optimal balance of diversity and difficulty. Comprehensive experiments conducted on both synthetic and real data justify the important theoretical properties and demonstrate the superior performance of BOSS as compared with the competitive baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- GORACS: Group-level Optimal Transport-guided Coreset Selection for LLM-based Recommender SystemsTiehua Mei, Hengrui Chen, Peng Yu, Jiaqing Liang et al.KDD 2025 · 3 citations
- Accelerating Benchmarking of Functional Connectivity Modeling via Structure-aware Core-set SelectionLing Zhan, Zhen Li, Junjie Huang, Tao JiaICLR 2026 · 1 citation
- Lightweight Dataset Pruning without Full Training via Example Difficulty and Prediction UncertaintyYeseul Cho, Baekrok Shin, Changmin Kang, Chulhee YunICML 2025
- Bandit Guided Submodular Curriculum for Adaptive Subset SelectionPrateek Chanda, Prayas Agrawal, Saral Sureka, Lokesh Reddy Polu et al.NeurIPS 2025
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
Related papers
- Coverage-centric Coreset Selection for High Pruning RatesHaizhong Zheng, Rui Liu, Fan Lai, Atul PrakashICLR 2023 · 3 citations
- Efficient Core-set Selection for Deep Learning Through Squared Loss MinimizationJianting ChenICML 2025
- Contributing Dimension Structure of Deep Feature for Coreset SelectionZhijing Wan, Zhixiang Wang, Yuran Wang, Zheng Wang et al.AAAI 2024 · 11 citations
- Training Data Subset Selection for Regression with Controlled Generalization ErrorDurga Sivasubramanian, Rishabh K. Iyer, Ganesh Ramakrishnan, Abir DeICML 2021 · 25 citations
- BWS: Best Window Selection Based on Sample Scores for Data Pruning across Broad RangesHoyong Choi, Nohyun Ki, Hye Won ChungICML 2024 · 9 citations
