Are Large-scale Soft Labels Necessary for Large-scale Dataset Distillation?
Lingao Xiao, Yang He
Abstract
In ImageNet-condensation, the storage for auxiliary soft labels exceeds that of the condensed dataset by over 30 times. However, are large-scale soft labels necessary for large-scale dataset distillation? In this paper, we first discover that the high within-class similarity in condensed datasets necessitates the use of large-scale soft labels. This high within-class similarity can be attributed to the fact that previous methods use samples from different classes to construct a single batch for batch normalization (BN) matching. To reduce the within-class similarity, we introduce class-wise supervision during the image synthesizing process by batching the samples within classes, instead of across classes. As a result, we can increase within-class diversity and reduce the size of required soft labels. A key benefit of improved image diversity is that soft label compression can be achieved through simple random pruning, eliminating the need for complex rule-based strategies. Experiments validate our discoveries. For example, when condensing ImageNet-1K to 200 images per class, our approach compresses the required soft labels from 113 GB to 2.8 GB (40x compression) with a 2.6% performance gain. Code is available at: https://github.com/he-y/soft-label-pruning-for-dataset-distillation
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 00efa597-0a17-4726-b5e4-0dda580f06efCited by top-tier papers11
- FADRM: Fast and Accurate Data Residual Matching for Dataset DistillationJiacheng Cui, Xinyue Bi, Yaxin Luo, Xiaohan Zhao et al.NeurIPS 2025 · 12 citations
- Unifying Dataset Pruning and Distillation for Efficient Large-scale CompressionLingao Xiao, Songhua Liu, Yang He, Xinchao WangICML 2026 · 6 citations
- DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion ModelsQichao Wang, Yunhong Lu, Hengyuan Cao, Junyi Zhang et al.CVPR 2026 · 4 citations
- FairDD: Fair Dataset DistillationQihang Zhou, Shenhao Fang, Shibo He, Wenchao Meng et al.NeurIPS 2025 · 3 citations
- Rectifying Soft-Label Entangled Bias in Long-Tailed Dataset DistillationChenyang Jiang, Hang Zhao, Xinyu Zhang, Zhengcen Li et al.NeurIPS 2025 · 1 citation
Builds on27
- EfficientNetV2: Smaller Models and Faster TrainingMingxing Tan, Quoc V. LeICML 2021 · 4,239 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
Related papers
- Heavy Labels Out! Dataset Distillation with Label Space LighteningRuonan Yu, Songhua Liu, Zigeng Chen, Jingwen Ye et al.ICCV 2025
- Beyond Soft Label: Dataset Distillation via Orthogonal Gradient MatchingDeyu Bo, Xinchao WangCVPR 2026
- A Label is Worth A Thousand Images in Dataset DistillationTian Qin, Zhiwei Deng, David Alvarez-MelisNeurIPS 2024 · 39 citations
- Diversity-Enhanced Distribution Alignment for Dataset DistillationHongcheng Li, Yucan Zhou, Xiaoyan Gu, Bo Li et al.ICCV 2025 · 1 citation
- Optimizing Distributional Geometry Alignment with Optimal Transport for Generative Dataset DistillationXiao Cui, Yulei Qin, Wengang Zhou, Hongsheng Li et al.NeurIPS 2025 · 5 citations
