BWS: Best Window Selection Based on Sample Scores for Data Pruning across Broad Ranges
Hoyong Choi, Nohyun Ki, Hye Won Chung
摘要
Data subset selection aims to find a smaller yet informative subset of a large dataset that can approximate the full-dataset training, addressing challenges associated with training neural networks on large-scale datasets. However, existing methods tend to specialize in either high or low selection ratio regimes, lacking a universal approach that consistently achieves competitive performance across a broad range of selection ratios. We introduce a universal and efficient data subset selection method, Best Window Selection (BWS), by proposing a method to choose the best window subset from samples ordered based on their difficulty scores. This approach offers flexibility by allowing the choice of window intervals that span from easy to difficult samples. Furthermore, we provide an efficient mechanism for selecting the best window subset by evaluating its quality using kernel ridge regression. Our experimental results demonstrate the superior performance of BWS compared to other baselines across a broad range of selection ratios over datasets, including CIFAR-10/100 and ImageNet, and the scenarios involving training from random initialization or fine-tuning of pre-trained models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- SNAP: Low-Latency Test-Time Adaptation with Sparse UpdatesHyeongheon Cha, Dong Min Kim, Hye Won Chung, Taesik Gong 等NeurIPS 2025 · 被引用 4 次
- Annotation-Efficient Coreset Selection for Context-dependent SegmentationJin Zhang, Zhe Cao, Biwen Yang, Ruiheng ZhangCVPR 2026
- Lightweight Dataset Pruning without Full Training via Example Difficulty and Prediction UncertaintyYeseul Cho, Baekrok Shin, Changmin Kang, Chulhee YunICML 2025
- Diffusion Reconstruction-based Data Likelihood Estimation for Core-Set SelectionMingyang Chen, Jiawei Du, Bo Huang, Yi Wang 等AAAI 2026
- ELFS: Label-Free Coreset Selection with Proxy Training DynamicsHaizhong Zheng, Elisa Tsai, Yifu Lu, Jiachen Sun 等ICLR 2025
它引用的顶会 Paper17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 被引用 806 次
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 被引用 784 次
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli 等NeurIPS 2022 · 被引用 720 次
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 被引用 674 次
相关 Paper
- Balancing Feature Similarity and Label Variability for Optimal Size-Aware One-shot Subset SelectionAbhinab Acharya, Dayou Yu, Qi Yu, Xumin LiuICML 2024 · 被引用 6 次
- Initializing Models with Larger OnesZhiqiu Xu, Yanjie Chen, Kirill Vishniakov, Yida Yin 等ICLR 2024 · 被引用 40 次
- Difficulty Is Not Enough: Curriculum Learning for LLMs Fine-tuning Must Consider UtilityZishang Jiang, Jinyi Han, Tingyun Li, Xinyi Wang 等AAAI 2026 · 被引用 1 次
- Training Data Subset Selection for Regression with Controlled Generalization ErrorDurga Sivasubramanian, Rishabh K. Iyer, Ganesh Ramakrishnan, Abir DeICML 2021 · 被引用 25 次
- Repeated Random Sampling for Minimizing the Time-to-Accuracy of LearningPatrik Okanovic, Roger Waleffe, Vasilis Mageirakos, Konstantinos E. Nikolakakis 等ICLR 2024 · 被引用 29 次
