OrderDP: A Theoretically Guaranteed Lossless Dynamic Data Pruning Framework
Chenhan Jin, Shengze Xu, Qingsong Wang, Fan JIA, Dingshuo Chen, Tieyong Zeng
摘要
Data pruning (DP), as an oft-stated strategy to alleviate heavy training burdens, reduces the volume of training samples according to a well-defined pruning method while striving for near-lossless performance. However, existing approaches, which commonly select highly informative samples, can lead to biased gradient estimation compared to full-dataset training. Furthermore, the analysis of this bias and its impact on final performance remains ambiguous. To address these challenges, we propose OrderDP, a plug-and-play framework that aims to obtain stable, unbiased, and near-lossless training acceleration with theoretical guarantees. Specifically, OrderDP first randomly selects a subset and then chooses the top-q samples, where unbiasedness is established with respect to a surrogate loss. This ensures that OrderDP conducts unbiased training in terms of the surrogate objective. We further establish convergence and generalization analyses, elucidating how OrderDP affects optimal performance and enables well-controlled acceleration while ensuring guaranteed final performance. Empirically, we evaluate OrderDP against comprehensive baselines on CIFAR-10, CIFAR-100, and ImageNet-1K, demonstrating competitive accuracy, stable convergence, and exact control-all with a simpler design and faster runtime, while reducing training cost by over 40%. Delivering both strong performance and computational efficiency, our method serves as a robust and easily adaptable tool for data-efficient learning. The code is publicly available at https://github.com/shengze-xu/OrderDP .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li 等AAAI 2020 · 被引用 4,134 次
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu 等ICLR 2020 · 被引用 1,170 次
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 被引用 806 次
相关 Paper
- Partial Forward Blocking: A Novel Data Pruning Paradigm for Lossless Training AccelerationDongyue Wu, Zilin Guo, Jialong Zuo, Nong Sang 等ICCV 2025 · 被引用 2 次
- Dataset Pruning: Reducing Training Data by Examining Generalization InfluenceShuo Yang, Zeke Xie, Hanyu Peng, Min Xu 等ICLR 2023 · 被引用 21 次
- InfoBatch: Lossless Training Speed Up by Unbiased Dynamic Data PruningZiheng Qin, Kai Wang, Zangwei Zheng, Jianyang Gu 等ICLR 2024 · 被引用 94 次
- Improving the Scaling Laws of Synthetic Data with Deliberate PracticeReyhane Askari Hemmat, Mohammad Pezeshki, Elvis Dohmatob, Florian Bordes 等ICML 2025
- Repeated Random Sampling for Minimizing the Time-to-Accuracy of LearningPatrik Okanovic, Roger Waleffe, Vasilis Mageirakos, Konstantinos E. Nikolakakis 等ICLR 2024 · 被引用 29 次
