Lightweight Dataset Pruning without Full Training via Example Difficulty and Prediction Uncertainty
Yeseul Cho, Baekrok Shin, Changmin Kang, Chulhee Yun
Abstract
Recent advances in deep learning rely heavily on massive datasets, leading to substantial storage and training costs. Dataset pruning aims to alleviate this demand by discarding redundant examples. However, many existing methods require training a model with a full dataset over a large number of epochs before being able to prune the dataset, which ironically makes the pruning process more expensive than just training the model on the entire dataset. To overcome this limitation, we introduce a Difficulty and Uncertainty-Aware Lightweight (DUAL) score, which aims to identify important samples from the early training stage by considering both example difficulty and prediction uncertainty. To address a catastrophic accuracy drop at extreme pruning, we further propose a ratio-adaptive sampling using Beta distribution. Experiments on various datasets and learning scenarios such as image classification with label noise, image corruption, and model architecture generalization demonstrate the superiority of our method over previous state-of-the-art (SOTA) approaches. Specifically, on ImageNet-1k, our method reduces the time cost for pruning to 66% compared to previous methods while achieving a SOTA, specifically 60% test accuracy at a 90% pruning ratio. On CIFAR datasets, the time cost is reduced to just 15% while maintaining SOTA performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8c261667-e21a-4d3f-9054-4401fdd9221cCited by top-tier papers8
- AdaSTaR: Adaptive Data Sampling for Training Self-Taught ReasonersReiss Koh, Wonbeen Oh, Jaein Jang, Minhyung Lee et al.NeurIPS 2025 · 8 citations
- Unifying Dataset Pruning and Distillation for Efficient Large-scale CompressionLingao Xiao, Songhua Liu, Yang He, Xinchao WangICML 2026 · 6 citations
- Rethinking Dataset Distillation: Hard Truths about Soft LabelsPriyam Dey, Aditya Sahdev, Sunny Bhati, Konda Reddy Mopuri et al.CVPR 2026 · 2 citations
- TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction TuningManish Nagaraj, Sakshi Choudhary, Utkarsh Saxena, Deepak Ravikumar et al.ICML 2026 · 2 citations
- Data Agent: Learning to Select Data via End-to-End Dynamic OptimizationSuorong Yang, Fangjian Su, Hai Gan, Ziqi Ye et al.ICML 2026
Builds on18
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman et al.ICLR 2020 · 462 citations
- Identifying Mislabeled Data using the Area Under the Margin RankingGeoff Pleiss, Tianyi Zhang, Ethan R. Elenberg, Kilian Q. WeinbergerNeurIPS 2020 · 398 citations
- Robust early-learning: Hindering the memorization of noisy labelsXiaobo Xia, Tongliang Liu, Bo Han, Chen Gong et al.ICLR 2021 · 322 citations
Related papers
- UNSEEN: Enhancing Dataset Pruning from a Generalization PerspectiveFurui Xu, Shaobo Wang, Jiajun Zhang, Chenghao Sun et al.AAAI 2026
- Spanning Training Progress: Temporal Dual-Depth Scoring (TDDS) for Enhanced Dataset PruningXin Zhang, Jiawei Du, Yunsong Li, Weiying Xie et al.CVPR 2024 · 12 citations
- DRoP: Distributionally Robust Data PruningArtem M. Vysogorets, Kartik Ahuja, Julia KempeICLR 2025
- Exploring Learning Complexity for Efficient Downstream Dataset PruningWenyu Jiang, Zhenlong Liu, Zejian Xie, Songxin Zhang et al.ICLR 2025
- Dataset Pruning: Reducing Training Data by Examining Generalization InfluenceShuo Yang, Zeke Xie, Hanyu Peng, Min Xu et al.ICLR 2023 · 21 citations
