Unifying Dataset Pruning and Distillation for Efficient Large-scale Compression
Lingao Xiao, Songhua Liu, Yang He, Xinchao Wang
摘要
Dataset pruning (DP) and dataset distillation (DD) fundamentally differ in their outputs: DP selects original image subsets, while DD generates synthetic images. Recently, DD's increasing reliance on original images suggests a convergence of the two directions. To investigate this convergence trend, we propose a unified dataset compression (DC) benchmark. This benchmark reveals an interesting trade-off for soft-label-DD: while soft labels provide valuable information, they can make the distillation process less essential, as distilled images may not always outperform random subsets. In addition, the benchmark reveals that in current stages, dataset pruning outperforms dataset distillation at small dataset sizes. Given these observations, we explore hard-label-DC as a complementary approach that emphasizes image quality while offering substantial storage efficiency. Our PCA (Prune, Combine, and Augment) is the first framework that does not rely on soft labels but instead focuses on image quality. (1) "P" means selecting easy samples based on dataset pruning metrics, (2) "C" indicates combining these samples effectively, and (3) "A" is to apply constrained image augmentation during training. Extensive experiments validate that PCA significantly outperforms existing DD and DP methods without soft labels. Code is at GitHub.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper45
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao 等CVPR 2022 · 被引用 2,138 次
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 被引用 806 次
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli 等NeurIPS 2022 · 被引用 720 次
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 被引用 684 次
相关 Paper
- A Label is Worth A Thousand Images in Dataset DistillationTian Qin, Zhiwei Deng, David Alvarez-MelisNeurIPS 2024 · 被引用 39 次
- Are Large-scale Soft Labels Necessary for Large-scale Dataset Distillation?Lingao Xiao, Yang HeNeurIPS 2024 · 被引用 19 次
- Rethinking Dataset Distillation: Hard Truths about Soft LabelsPriyam Dey, Aditya Sahdev, Sunny Bhati, Konda Reddy Mopuri 等CVPR 2026 · 被引用 2 次
- Post Training Quantization for Efficient Dataset CondensationLinh-Tam Tran, Sung-Ho BaeAAAI 2026
- Dataset QuantizationDaquan Zhou, Kai Wang, Jianyang Gu, Xiangyu Peng 等ICCV 2023 · 被引用 65 次
