Dataset Pruning: Reducing Training Data by Examining Generalization Influence
Shuo Yang, Zeke Xie, Hanyu Peng, Min Xu, Mingming Sun, Ping Li
摘要
The great success of deep learning heavily relies on increasingly larger training data, which comes at a price of huge computational and infrastructural costs. This poses crucial questions that, do all training data contribute to model's performance? How much does each individual training sample or a sub-training-set affect the model's generalization, and how to construct the smallest subset from the entire training data as a proxy training set without significantly sacrificing the model's performance? To answer these, we propose dataset pruning, an optimization-based sample selection method that can (1) examine the influence of removing a particular set of training samples on model's generalization ability with theoretical guarantee, and (2) construct the smallest subset of training data that yields strictly constrained generalization gap. The empirically observed generalization gap of dataset pruning is substantially consistent with our theoretical expectations. Furthermore, the proposed method prunes 40% training examples on the CIFAR-10 dataset, halves the convergence time with only 1.3% test accuracy decrease, which is superior to previous score-based sample selection methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper72
- Model Sparsity Can Simplify Machine UnlearningJinghan Jia, Jiancheng Liu, Parikshit Ram, Yuguang Yao 等NeurIPS 2023 · 被引用 293 次
- Data-efficient Fine-tuning for LLM-based RecommendationXinyu Lin, Wenjie Wang, Yongqi Li, Shuo Yang 等SIGIR 2024 · 被引用 152 次
- InfoBatch: Lossless Training Speed Up by Unbiased Dynamic Data PruningZiheng Qin, Kai Wang, Zangwei Zheng, Jianyang Gu 等ICLR 2024 · 被引用 94 次
- Data Pruning via Moving-one-Sample-outHaoru Tan, Sitong Wu, Fei Du, Yukang Chen 等NeurIPS 2023 · 被引用 91 次
- Condensing Graphs via One-Step Gradient MatchingWei Jin, Xianfeng Tang, Haoming Jiang, Zheng Li 等KDD 2022 · 被引用 68 次
它引用的顶会 Paper12
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 被引用 806 次
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 被引用 390 次
- Free Lunch for Few-shot Learning: Distribution CalibrationShuo Yang, Lu Liu, Min XuICLR 2021 · 被引用 378 次
- Coresets via Bilevel Optimization for Continual Learning and StreamingZalán Borsos, Mojmir Mutny, Andreas KrauseNeurIPS 2020 · 被引用 320 次
- Dataset Distillation with Infinitely Wide Convolutional NetworksTimothy Nguyen, Roman Novak, Lechao Xiao, Jaehoon LeeNeurIPS 2021 · 被引用 313 次
相关 Paper
- UNSEEN: Enhancing Dataset Pruning from a Generalization PerspectiveFurui Xu, Shaobo Wang, Jiajun Zhang, Chenghao Sun 等AAAI 2026
- OrderDP: A Theoretically Guaranteed Lossless Dynamic Data Pruning FrameworkChenhan Jin, Shengze Xu, Qingsong Wang, Fan JIA 等ICLR 2026 · 被引用 3 次
- Spanning Training Progress: Temporal Dual-Depth Scoring (TDDS) for Enhanced Dataset PruningXin Zhang, Jiawei Du, Yunsong Li, Weiying Xie 等CVPR 2024 · 被引用 12 次
- DynaMS: Dyanmic Margin Selection for Efficient Deep LearningJiaxing Wang, Yong Li, Jingwei Zhuo, Xupeng Shi 等ICLR 2023
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman 等ICLR 2020 · 被引用 462 次
