MGDD: A Meta Generator for Fast Dataset Distillation
Songhua Liu, Xinchao Wang
摘要
Existing dataset distillation (DD) techniques typically rely on iterative strategies to synthesize condensed datasets, where datasets before and after distillation are forward and backward through neural networks a massive number of times. Despite the promising results achieved, the time efficiency of prior approaches is still far from satisfactory. Moreover, when different sizes of synthetic datasets are required, they have to repeat the iterative training procedures, which is highly cumbersome and lacks flexibility. In this paper, different from the time-consuming forward-backward passes, we introduce a generative fashion for dataset distillation with significantly improved efficiency. Specifically, synthetic samples are produced by a generator network conditioned on the initialization of DD, while synthetic labels are obtained by solving a least-squares problem in a feature space. Our theoretical analysis reveals that the errors of synthetic datasets solved in the original space and then processed by any conditional generators are upper-bounded. To find a satisfactory generator efficiently, we propose a meta-learning algorithm, where a meta generator is trained on a large dataset so that only a few steps are required to adapt to a target dataset. The meta generator is termed as MGDD in our approach. Once adapted, it can handle arbitrary sizes of synthetic datasets, even for those unseen during adaptation. Experiments demonstrate that the generator adapted with only a limited number of steps performs on par with those state-of-the-art DD methods and yields 22 × acceleration.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Diversity-Driven Synthesis: Enhancing Dataset Distillation through Directed Weight AdjustmentJiawei Du, Xin Zhang, Juncheng Hu, Wenxin Huang 等NeurIPS 2024 · 被引用 43 次
- Ungeneralizable ExamplesJingwen Ye, Xinchao WangCVPR 2024 · 被引用 3 次
- Understanding Dataset Distillation via Spectral FilteringDeyu Bo, Songhua Liu, Xinchao WangICLR 2026 · 被引用 3 次
- Distilled Datamodel with Reverse Gradient MatchingJingwen Ye, Ruonan Yu, Songhua Liu, Xinchao WangCVPR 2024 · 被引用 1 次
- Mind Your Margin and Boundary: Are Your Distilled Datasets Truly Robust?Muquan Li, Yingyi Ma, Yihong Huang, Hang Gou 等ICML 2026
它引用的顶会 Paper24
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 被引用 5,234 次
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 被引用 684 次
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 被引用 390 次
相关 Paper
- Accelerating Dataset Distillation via Model AugmentationLei Zhang, Jie Zhang, Bowen Lei, Subhabrata Mukherjee 等CVPR 2023
- Data Distillation Can Be Like Vodka: Distilling More Times For Better QualityXuxi Chen, Yu Yang, Zhangyang Wang, Baharan MirzasoleimanICLR 2024 · 被引用 19 次
- Few-Shot Dataset Distillation via Translative Pre-TrainingSonghua Liu, Xinchao WangICCV 2023 · 被引用 15 次
- TGDD: Trajectory Guided Dataset Distillation with Balanced DistributionFengli Ran, Xiao Pu, Bo Liu, Xiuli Bi 等AAAI 2026
- DREAM: Efficient Dataset Distillation by Representative MatchingYanqing Liu, Jianyang Gu, Kai Wang, Zheng Zhu 等ICCV 2023 · 被引用 114 次
