Data Distillation Can Be Like Vodka: Distilling More Times For Better Quality
Xuxi Chen, Yu Yang, Zhangyang Wang, Baharan Mirzasoleiman
Abstract
Dataset distillation aims to minimize the time and memory needed for training deep networks on large datasets, by creating a small set of synthetic images that has a similar generalization performance to that of the full dataset. However, current dataset distillation techniques fall short, showing a notable performance gap when compared to training on the original data. In this work, we are the first to argue that using just one synthetic subset for distillation will not yield optimal generalization performance. This is because the training dynamics of deep networks drastically change during the training. Hence, multiple synthetic subsets are required to capture the training dynamics at different phases of training. To address this issue, we propose Progressive Dataset Distillation (PDD). PDD synthesizes multiple small sets of synthetic images, each conditioned on the previous sets, and trains the model on the cumulative union of these subsets without requiring additional training time. Our extensive experiments show that PDD can effectively improve the performance of existing dataset distillation methods by up to 4.3%. In addition, our method for the first time enable generating considerably larger synthetic datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2930de0-a13e-4eab-b61c-0bef88ac9b37Cited by top-tier papers13
- Towards Lossless Dataset Distillation via Difficulty-Aligned Trajectory MatchingZiyao Guo, Kai Wang, George Cazenavette, Hui Li et al.ICLR 2024 · 142 citations
- Out-of-Distribution Detection with Relative AnglesBerker Demirel, Marco Fumero, Francesco LocatelloNeurIPS 2025 · 3 citations
- FairDD: Fair Dataset DistillationQihang Zhou, Shenhao Fang, Shibo He, Wenchao Meng et al.NeurIPS 2025 · 3 citations
- Multimodal Dataset Distillation via Phased Teacher ModelsShengbin Guo, Hang Zhao, Senqiao Yang, Chenyang Jiang et al.ICLR 2026 · 1 citation
- SCAN: Bootstrapping Contrastive Pre-training for Data EfficiencyYangyang Guo, Mohan KankanhalliICCV 2025 · 1 citation
Builds on17
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- The Pitfalls of Simplicity Bias in Neural NetworksHarshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain et al.NeurIPS 2020 · 503 citations
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 390 citations
- Dataset Distillation with Infinitely Wide Convolutional NetworksTimothy Nguyen, Roman Novak, Lechao Xiao, Jaehoon LeeNeurIPS 2021 · 313 citations
- Dataset Meta-Learning from Kernel Ridge-RegressionTimothy Nguyen, Zhourong Chen, Jaehoon LeeICLR 2021 · 307 citations
Related papers
- Accelerating Dataset Distillation via Model AugmentationLei Zhang, Jie Zhang, Bowen Lei, Subhabrata Mukherjee et al.CVPR 2023
- Sequential Subset Matching for Dataset DistillationJiawei Du, Qin Shi, Joey Tianyi ZhouNeurIPS 2023 · 52 citations
- Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-training of Deep NetworksSiddharth Joshi, Jiayi Ni, Baharan MirzasoleimanICLR 2025
- MGDD: A Meta Generator for Fast Dataset DistillationSonghua Liu, Xinchao WangNeurIPS 2023 · 13 citations
- TGDD: Trajectory Guided Dataset Distillation with Balanced DistributionFengli Ran, Xiao Pu, Bo Liu, Xiuli Bi et al.AAAI 2026
