Taming Diffusion for Dataset Distillation with High Representativeness
Lin Zhao, Yushu Wu, Xinru Jiang, Jianyang Gu, Yanzhi Wang, Xiaolin Xu, Pu Zhao, Xue Lin
摘要
Recent deep learning models demand larger datasets, driving the need for dataset distillation to create compact, cost-efficient datasets while maintaining performance. Due to the powerful image generation capability of diffusion, it has been introduced to this field for generating distilled images. In this paper, we systematically investigate issues present in current diffusion-based dataset distillation methods, including inaccurate distribution matching, distribution deviation with random noise, and separate sampling. Building on this, we propose D 3 HR, a novel diffusionbased framework to generate distilled datasets with high representativeness. Specifically, we adopt DDIM inversion to map the latents of the full dataset from a low-normality latent domain to a high-normality Gaussian domain, preserving information and ensuring structural consistency to generate representative latents for the distilled dataset. Furthermore, we propose an efficient sampling scheme to better align the representative latents with the high-normality Gaussian distribution. Our comprehensive experiments demonstrate that D 3 HR can achieve higher accuracy across different model architectures compared with state-of-the-art baselines in dataset distillation. Source code: https://github. com/lin-zhao-resoLve/D3HR .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- HierAmp: Coarse-to-Fine Autoregressive Amplification for Generative Dataset DistillationLin Zhao, Xinru Jiang, Xi Xiao, Qihui Fan 等CVPR 2026 · 被引用 9 次
- Not All Directions Matter: Towards Structured and Task-Aware Low-Rank Model AdaptationXi Xiao, Chenrui Ma, Yunbei Zhang, Chen Liu 等ACL 2026 · 被引用 6 次
- Prime Once, then Reprogram Locally: An Efficient Alternative to Black-Box Service Model AdaptationYunbei Zhang, Chengyi Cai, Feng Liu, Jihun HammCVPR 2026 · 被引用 5 次
- Diffusion Models as Dataset Distillation PriorsDuo Su, Huyu Wu, Huanran Chen, Yiming Shi 等ICLR 2026 · 被引用 3 次
- IMS3: Breaking Distributional Aggregation in Diffusion-Based Dataset DistillationChenru Wang, Yunyi Chen, Zijun Yang, Joey Tianyi Zhou 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper28
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- An Adaptive Sampling Framework for Diffusion-based Dataset Distillation with High Fidelity and DiversitySunbeom Jeong, Sehwan Kim, Hyeonggeun Han, Hyungjun Joo 等AAAI 2026
- Unlocking Dataset Distillation with Diffusion ModelsBrian B. Moser, Federico Raue, Sebastian Palacio, Stanislav Frolov 等NeurIPS 2025 · 被引用 23 次
- Noise-Optimized Distribution Distillation for Dataset CondensationTongfei Liu, Yufan Liu, Bing Li, Weiming Hu 等ACM MM 2025
- Influence-Guided Diffusion for Dataset DistillationMingyang Chen, Jiawei Du, Bo Huang, Yi Wang 等ICLR 2025
- D4M: Dataset Distillation via Disentangled Diffusion ModelDuo Su, Junjie Hou, Weizhi Gao, Yingjie Tian 等CVPR 2024 · 被引用 11 次
