BWS: Best Window Selection Based on Sample Scores for Data Pruning across Broad Ranges
Hoyong Choi, Nohyun Ki, Hye Won Chung
Abstract
Data subset selection aims to find a smaller yet informative subset of a large dataset that can approximate the full-dataset training, addressing challenges associated with training neural networks on large-scale datasets. However, existing methods tend to specialize in either high or low selection ratio regimes, lacking a universal approach that consistently achieves competitive performance across a broad range of selection ratios. We introduce a universal and efficient data subset selection method, Best Window Selection (BWS), by proposing a method to choose the best window subset from samples ordered based on their difficulty scores. This approach offers flexibility by allowing the choice of window intervals that span from easy to difficult samples. Furthermore, we provide an efficient mechanism for selecting the best window subset by evaluating its quality using kernel ridge regression. Our experimental results demonstrate the superior performance of BWS compared to other baselines across a broad range of selection ratios over datasets, including CIFAR-10/100 and ImageNet, and the scenarios involving training from random initialization or fine-tuning of pre-trained models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7bb35cbb-31d9-476d-a07c-15ea73181460Cited by top-tier papers5
- SNAP: Low-Latency Test-Time Adaptation with Sparse UpdatesHyeongheon Cha, Dong Min Kim, Hye Won Chung, Taesik Gong et al.NeurIPS 2025 · 4 citations
- Annotation-Efficient Coreset Selection for Context-dependent SegmentationJin Zhang, Zhe Cao, Biwen Yang, Ruiheng ZhangCVPR 2026
- Lightweight Dataset Pruning without Full Training via Example Difficulty and Prediction UncertaintyYeseul Cho, Baekrok Shin, Changmin Kang, Chulhee YunICML 2025
- Diffusion Reconstruction-based Data Likelihood Estimation for Core-Set SelectionMingyang Chen, Jiawei Du, Bo Huang, Yi Wang et al.AAAI 2026
- ELFS: Label-Free Coreset Selection with Proxy Training DynamicsHaizhong Zheng, Elisa Tsai, Yifu Lu, Jiachen Sun et al.ICLR 2025
Builds on17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 674 citations
Related papers
- Balancing Feature Similarity and Label Variability for Optimal Size-Aware One-shot Subset SelectionAbhinab Acharya, Dayou Yu, Qi Yu, Xumin LiuICML 2024 · 6 citations
- Initializing Models with Larger OnesZhiqiu Xu, Yanjie Chen, Kirill Vishniakov, Yida Yin et al.ICLR 2024 · 40 citations
- Difficulty Is Not Enough: Curriculum Learning for LLMs Fine-tuning Must Consider UtilityZishang Jiang, Jinyi Han, Tingyun Li, Xinyi Wang et al.AAAI 2026 · 1 citation
- Training Data Subset Selection for Regression with Controlled Generalization ErrorDurga Sivasubramanian, Rishabh K. Iyer, Ganesh Ramakrishnan, Abir DeICML 2021 · 25 citations
- Repeated Random Sampling for Minimizing the Time-to-Accuracy of LearningPatrik Okanovic, Roger Waleffe, Vasilis Mageirakos, Konstantinos E. Nikolakakis et al.ICLR 2024 · 29 citations
