Lightweight Dataset Pruning without Full Training via Example Difficulty and Prediction Uncertainty
Yeseul Cho, Baekrok Shin, Changmin Kang, Chulhee Yun
摘要
Recent advances in deep learning rely heavily on massive datasets, leading to substantial storage and training costs. Dataset pruning aims to alleviate this demand by discarding redundant examples. However, many existing methods require training a model with a full dataset over a large number of epochs before being able to prune the dataset, which ironically makes the pruning process more expensive than just training the model on the entire dataset. To overcome this limitation, we introduce a Difficulty and Uncertainty-Aware Lightweight (DUAL) score, which aims to identify important samples from the early training stage by considering both example difficulty and prediction uncertainty. To address a catastrophic accuracy drop at extreme pruning, we further propose a ratio-adaptive sampling using Beta distribution. Experiments on various datasets and learning scenarios such as image classification with label noise, image corruption, and model architecture generalization demonstrate the superiority of our method over previous state-of-the-art (SOTA) approaches. Specifically, on ImageNet-1k, our method reduces the time cost for pruning to 66% compared to previous methods while achieving a SOTA, specifically 60% test accuracy at a 90% pruning ratio. On CIFAR datasets, the time cost is reduced to just 15% while maintaining SOTA performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- AdaSTaR: Adaptive Data Sampling for Training Self-Taught ReasonersReiss Koh, Wonbeen Oh, Jaein Jang, Minhyung Lee 等NeurIPS 2025 · 被引用 8 次
- Unifying Dataset Pruning and Distillation for Efficient Large-scale CompressionLingao Xiao, Songhua Liu, Yang He, Xinchao WangICML 2026 · 被引用 6 次
- Rethinking Dataset Distillation: Hard Truths about Soft LabelsPriyam Dey, Aditya Sahdev, Sunny Bhati, Konda Reddy Mopuri 等CVPR 2026 · 被引用 2 次
- TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction TuningManish Nagaraj, Sakshi Choudhary, Utkarsh Saxena, Deepak Ravikumar 等ICML 2026 · 被引用 2 次
- Data Agent: Learning to Select Data via End-to-End Dynamic OptimizationSuorong Yang, Fangjian Su, Hai Gan, Ziqi Ye 等ICML 2026
它引用的顶会 Paper18
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 被引用 806 次
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli 等NeurIPS 2022 · 被引用 720 次
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman 等ICLR 2020 · 被引用 462 次
- Identifying Mislabeled Data using the Area Under the Margin RankingGeoff Pleiss, Tianyi Zhang, Ethan R. Elenberg, Kilian Q. WeinbergerNeurIPS 2020 · 被引用 398 次
- Robust early-learning: Hindering the memorization of noisy labelsXiaobo Xia, Tongliang Liu, Bo Han, Chen Gong 等ICLR 2021 · 被引用 322 次
相关 Paper
- UNSEEN: Enhancing Dataset Pruning from a Generalization PerspectiveFurui Xu, Shaobo Wang, Jiajun Zhang, Chenghao Sun 等AAAI 2026
- Spanning Training Progress: Temporal Dual-Depth Scoring (TDDS) for Enhanced Dataset PruningXin Zhang, Jiawei Du, Yunsong Li, Weiying Xie 等CVPR 2024 · 被引用 12 次
- DRoP: Distributionally Robust Data PruningArtem M. Vysogorets, Kartik Ahuja, Julia KempeICLR 2025
- Exploring Learning Complexity for Efficient Downstream Dataset PruningWenyu Jiang, Zhenlong Liu, Zejian Xie, Songxin Zhang 等ICLR 2025
- Dataset Pruning: Reducing Training Data by Examining Generalization InfluenceShuo Yang, Zeke Xie, Hanyu Peng, Min Xu 等ICLR 2023 · 被引用 21 次
