Noise-Optimized Distribution Distillation for Dataset Condensation
Tongfei Liu, Yufan Liu, Bing Li, Weiming Hu, Yuming Li, Chenguang Ma
摘要
Dataset condensation distills a large dataset into a small synthetic surrogate dataset with similar training efficacy on downstream tasks. Of the existing condensation methods, diffusion-based methods that synthesize surrogate datasets with diffusion models have successfully distilled high-resolution datasets with high training efficacy and satisfactory cross-architectural transferability. However, these methods exhibit a random sampling bias that impairs their performance in dataset condensation settings. We propose a novel dataset condensation method called Noise-Optimized Distribution Distillation (NODD) that mitigates this sampling bias to improve the training performance of synthetic datasets generated with diffusion models. NODD can integrate with existing diffusion-based methods to produce synthetic datasets with enhanced training performance.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- An Adaptive Sampling Framework for Diffusion-based Dataset Distillation with High Fidelity and DiversitySunbeom Jeong, Sehwan Kim, Hyeonggeun Han, Hyungjun Joo 等AAAI 2026
- Taming Diffusion for Dataset Distillation with High RepresentativenessLin Zhao, Yushu Wu, Xinru Jiang, Jianyang Gu 等ICML 2025
- CaO2: Rectifying Inconsistencies in Diffusion-Based Dataset DistillationHaoxuan Wang, Zhenghao Zhao, Junyi Wu, Yuzhang Shang 等ICCV 2025 · 被引用 1 次
- Unlocking Dataset Distillation with Diffusion ModelsBrian B. Moser, Federico Raue, Sebastian Palacio, Stanislav Frolov 等NeurIPS 2025 · 被引用 23 次
- Efficient Dataset Distillation via Minimax DiffusionJianyang Gu, Saeed Vahidian, Vyacheslav Kungurtsev, Haonan Wang 等CVPR 2024
