A Label is Worth A Thousand Images in Dataset Distillation
Tian Qin, Zhiwei Deng, David Alvarez-Melis
摘要
Data is a crucial factor in the performance of machine learning models, a principle that dataset distillation methods exploit by compressing training datasets into much smaller counterparts that maintain similar downstream performance. Understanding how and why data distillation methods work is vital not only for improving these methods but also for revealing fundamental characteristics of"good"training data. However, a major challenge in achieving this goal is the observation that distillation approaches, which rely on sophisticated but mostly disparate methods to generate synthetic data, have little in common with each other. In this work, we highlight a largely overlooked aspect common to most of these methods: the use of soft (probabilistic) labels. Through a series of ablation experiments, we study the role of soft labels in depth. Our results reveal that the main factor explaining the performance of state-of-the-art distillation methods is not the specific techniques used to generate synthetic data but rather the use of soft labels. Furthermore, we demonstrate that not all soft labels are created equal; they must contain to be beneficial. We also provide empirical scaling laws that characterize the effectiveness of soft labels as a function of images-per-class in the distilled dataset and establish an empirical Pareto frontier for data-efficient learning. Combined, our findings challenge conventional wisdom in dataset distillation, underscore the importance of soft labels in learning, and suggest new directions for improving distillation methods. Code for all experiments is available at https://github.com/sunnytqin/no-distillation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Dataset Distillation for Memorized Data: Soft Labels can Leak Held-Out Teacher KnowledgeFreya Behrens, Lenka ZdeborováICLR 2026 · 被引用 9 次
- Unifying Dataset Pruning and Distillation for Efficient Large-scale CompressionLingao Xiao, Songhua Liu, Yang He, Xinchao WangICML 2026 · 被引用 6 次
- Grounding and Enhancing Informativeness and Utility in Dataset DistillationShaobo Wang, Yantai Yang, Guo Chen, Peiru Li 等ICLR 2026 · 被引用 3 次
- Rectified Decoupled Dataset Distillation: A Closer Look for Fair and Comprehensive EvaluationXinhao Zhong, Shuoyang Sun, Xulin Gu, Chenyang Zhu 等ICLR 2026 · 被引用 2 次
- Rethinking Dataset Distillation: Hard Truths about Soft LabelsPriyam Dey, Aditya Sahdev, Sunny Bhati, Konda Reddy Mopuri 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper19
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 被引用 684 次
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 被引用 390 次
- Dataset Meta-Learning from Kernel Ridge-RegressionTimothy Nguyen, Zhourong Chen, Jaehoon LeeICLR 2021 · 被引用 307 次
- Dataset Distillation using Neural Feature RegressionYongchao Zhou, Ehsan Nezhadarya, Jimmy BaNeurIPS 2022 · 被引用 234 次
相关 Paper
- Going Beyond Feature Similarity: Effective Dataset distillation based on Class-aware Conditional Mutual InformationXinhao Zhong, Bin Chen, Hao Fang, Xulin Gu 等ICLR 2025
- Are Large-scale Soft Labels Necessary for Large-scale Dataset Distillation?Lingao Xiao, Yang HeNeurIPS 2024 · 被引用 19 次
- Enhancing Dataset Distillation via Non-Critical Region RefinementMinh-Tuan Tran, Trung Le, Xuan-May Le, Thanh-Toan Do 等CVPR 2025
- Hard Labels In! Rethinking the Role of Hard Labels in Mitigating Local Semantic DriftJiacheng Cui, Bingkui Tong, Xinyue Bi, Xiaohan Zhao 等ICML 2026
- Heavy Labels Out! Dataset Distillation with Label Space LighteningRuonan Yu, Songhua Liu, Zigeng Chen, Jingwen Ye 等ICCV 2025
