A Label is Worth A Thousand Images in Dataset Distillation
Tian Qin, Zhiwei Deng, David Alvarez-Melis
Abstract
Data is a crucial factor in the performance of machine learning models, a principle that dataset distillation methods exploit by compressing training datasets into much smaller counterparts that maintain similar downstream performance. Understanding how and why data distillation methods work is vital not only for improving these methods but also for revealing fundamental characteristics of"good"training data. However, a major challenge in achieving this goal is the observation that distillation approaches, which rely on sophisticated but mostly disparate methods to generate synthetic data, have little in common with each other. In this work, we highlight a largely overlooked aspect common to most of these methods: the use of soft (probabilistic) labels. Through a series of ablation experiments, we study the role of soft labels in depth. Our results reveal that the main factor explaining the performance of state-of-the-art distillation methods is not the specific techniques used to generate synthetic data but rather the use of soft labels. Furthermore, we demonstrate that not all soft labels are created equal; they must contain to be beneficial. We also provide empirical scaling laws that characterize the effectiveness of soft labels as a function of images-per-class in the distilled dataset and establish an empirical Pareto frontier for data-efficient learning. Combined, our findings challenge conventional wisdom in dataset distillation, underscore the importance of soft labels in learning, and suggest new directions for improving distillation methods. Code for all experiments is available at https://github.com/sunnytqin/no-distillation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 575512ae-a6a2-4888-96df-d8def7e81707Cited by top-tier papers11
- Dataset Distillation for Memorized Data: Soft Labels can Leak Held-Out Teacher KnowledgeFreya Behrens, Lenka ZdeborováICLR 2026 · 9 citations
- Unifying Dataset Pruning and Distillation for Efficient Large-scale CompressionLingao Xiao, Songhua Liu, Yang He, Xinchao WangICML 2026 · 6 citations
- Grounding and Enhancing Informativeness and Utility in Dataset DistillationShaobo Wang, Yantai Yang, Guo Chen, Peiru Li et al.ICLR 2026 · 3 citations
- Rectified Decoupled Dataset Distillation: A Closer Look for Fair and Comprehensive EvaluationXinhao Zhong, Shuoyang Sun, Xulin Gu, Chenyang Zhu et al.ICLR 2026 · 2 citations
- Rethinking Dataset Distillation: Hard Truths about Soft LabelsPriyam Dey, Aditya Sahdev, Sunny Bhati, Konda Reddy Mopuri et al.CVPR 2026 · 2 citations
Builds on19
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 390 citations
- Dataset Meta-Learning from Kernel Ridge-RegressionTimothy Nguyen, Zhourong Chen, Jaehoon LeeICLR 2021 · 307 citations
- Dataset Distillation using Neural Feature RegressionYongchao Zhou, Ehsan Nezhadarya, Jimmy BaNeurIPS 2022 · 234 citations
Related papers
- Going Beyond Feature Similarity: Effective Dataset distillation based on Class-aware Conditional Mutual InformationXinhao Zhong, Bin Chen, Hao Fang, Xulin Gu et al.ICLR 2025
- Are Large-scale Soft Labels Necessary for Large-scale Dataset Distillation?Lingao Xiao, Yang HeNeurIPS 2024 · 19 citations
- Enhancing Dataset Distillation via Non-Critical Region RefinementMinh-Tuan Tran, Trung Le, Xuan-May Le, Thanh-Toan Do et al.CVPR 2025
- Hard Labels In! Rethinking the Role of Hard Labels in Mitigating Local Semantic DriftJiacheng Cui, Bingkui Tong, Xinyue Bi, Xiaohan Zhao et al.ICML 2026
- Heavy Labels Out! Dataset Distillation with Label Space LighteningRuonan Yu, Songhua Liu, Zigeng Chen, Jingwen Ye et al.ICCV 2025
