Stochastic Forward-Backward Deconvolution: Training Diffusion Models with Finite Noisy Datasets
Haoye Lu, Qifan Wu, Yaoliang Yu
Abstract
Recent diffusion-based generative models achieve remarkable results by training on massive datasets, yet this practice raises concerns about memorization and copyright infringement. A proposed remedy is to train exclusively on noisy data with potential copyright issues, ensuring the model never observes original content. However, through the lens of deconvolution theory, we show that although it is theoretically feasible to learn the data distribution from noisy samples, the practical challenge of collecting sufficient samples makes successful learning nearly unattainable. To overcome this limitation, we propose to pretrain the model with a small fraction of clean data to guide the deconvolution process. Combined with our Stochastic Forward-Backward Deconvolution (SFBD) method, we attain FID 6.31 on CIFAR-10 with just 4% clean images (and 3.58 with 10%). We theoretically show that SFBD guides the model to learn the true data distribution. The result also highlights the importance of pretraining on limited but clean data or the alternative from similar datasets. Empirical studies further support these findings and offer additional insights.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Generative Modeling from Black-Box Corruptions via Self-Consistent Stochastic InterpolantsChirag Modi, Jiequn Han, Eric Vanden-Eijnden, Joan BrunaICLR 2026 · 4 citations
- Ambient Dataloops: Generative Models for Dataset RefinementAdrian Rodriguez-Munoz, William Daspit, Adam Klivans, Antonio Torralba et al.ICML 2026
- SFBD-OMNI: Bridge models for lossy measurement restoration with limited clean samplesHaoye Lu, Yaoliang Yu, Darren LoICLR 2026
- MAD: Manifold Attracted DiffusionDennis Elbrächter, Giovanni S. Alberti, Matteo SantacesariaICML 2026
Builds on22
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- Does Generation Require Memorization? Creative Diffusion Models using Ambient DiffusionKulin Shah, Alkis Kalavasis, Adam R. Klivans, Giannis DarasICML 2025
- Ambient Diffusion: Learning Clean Distributions from Corrupted DataGiannis Daras, Kulin Shah, Yuval Dagan, Aravind Gollakota et al.NeurIPS 2023 · 141 citations
- Forward-Learned Discrete Diffusion: Learning how to noise to denoise fasterGrigory Bartosh, Teodora Pandeva, Sushrut Karmalkar, Javier ZazoICLR 2026 · 4 citations
- Truncated Diffusion Probabilistic Models and Diffusion-based Adversarial Auto-EncodersHuangjie Zheng, Pengcheng He, Weizhu Chen, Mingyuan ZhouICLR 2023 · 19 citations
- Self-diffusion for Solving Inverse ProblemsGuanxiong Luo, Shoujin HuangNeurIPS 2025 · 5 citations
