Slight Corruption in Pre-training Data Makes Better Diffusion Models
Hao Chen, Yujin Han, Diganta Misra, Xiang Li, Kai Hu, Difan Zou, Masashi Sugiyama, Jindong Wang, Bhiksha Raj
Abstract
Diffusion models (DMs) have shown remarkable capabilities in generating realistic high-quality images, audios, and videos. They benefit significantly from extensive pre-training on large-scale datasets, including web-crawled data with paired data and conditions, such as image-text and image-class pairs. Despite rigorous filtering, these pre-training datasets often inevitably contain corrupted pairs where conditions do not accurately describe the data. This paper presents the first comprehensive study on the impact of such corruption in pre-training data of DMs. We synthetically corrupt ImageNet-1K and CC3M to pre-train and evaluate over 50 conditional DMs. Our empirical findings reveal that various types of slight corruption in pre-training can significantly enhance the quality, diversity, and fidelity of the generated images across different DMs, both during pre-training and downstream adaptation stages. Theoretically, we consider a Gaussian mixture model and prove that slight corruption in the condition leads to higher entropy and a reduced 2-Wasserstein distance to the ground truth of the data distribution generated by the corruptly trained DMs. Inspired by our analysis, we propose a simple method to improve the training of DMs on practical datasets by adding condition embedding perturbations (CEP). CEP significantly improves the performance of various DMs in both pre-training and downstream tasks. We hope that our study provides new insights into understanding the data and pre-training processes of DMs and all models are released at https://huggingface.co/DiffusionNoise.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 735d1d87-901c-43c7-81f7-72c96198e2c2Cited by top-tier papers6
- Golden Noise for Diffusion Models: A Learning FrameworkZikai Zhou, Shitong Shao, Lichen Bai, Shufei Zhang et al.ICCV 2025 · 9 citations
- Beyond the Golden Data: Resolving the Motion-Vision Quality Dilemma via Timestep Selective TrainingXiangyang Luo, Qingyu Li, Yuming Li, Guanbo Huang et al.CVPR 2026 · 3 citations
- Guiding Noisy Label Conditional Diffusion Models with Score-Based Discriminator CorrectionNguyen Cong Dat, Bao Hieu Tran, Tung Hoang-ThanhICCV 2025 · 3 citations
- Bidirectional Noise Injection: Enhancing Diffusion Models via Coordinated Input-Output PerturbationTianyi Zheng, Jiayang Gao, Peng-Tao Jiang, Fengxiang Yang et al.AAAI 2026
- Personalized Federated Training of Diffusion Models with Privacy GuaranteesKumar Kshitij Patel, Bingqing Jiang, A. F. M. Mahfuzul Kabir, Weitong Zhang et al.CVPR 2026
Builds on67
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Ambient Diffusion Omni: Training Good Models with Bad DataGiannis Daras, Adrián Rodríguez-Muñoz, Adam R. Klivans, Antonio Torralba et al.NeurIPS 2025 · 17 citations
- How Much is a Noisy Image Worth? Data Scaling Laws for Ambient DiffusionGiannis Daras, Yeshwanth Cherapanamjeri, Constantinos DaskalakisICLR 2025
- One Transformer Fits All Distributions in Multi-Modal Diffusion at ScaleFan Bao, Shen Nie, Kaiwen Xue, Chongxuan Li et al.ICML 2023 · 236 citations
- Diffusion Adaptive Text Embedding for Text-to-Image Diffusion ModelsByeonghu Na, Minsang Park, Gyuwon Sim, Donghyeok Shin et al.NeurIPS 2025 · 8 citations
- Diffusion Models as Masked AutoencodersChen Wei, Karttikeya Mangalam, Po-Yao Huang, Yanghao Li et al.ICCV 2023 · 82 citations
