Stochastic Forward-Backward Deconvolution: Training Diffusion Models with Finite Noisy Datasets
Haoye Lu, Qifan Wu, Yaoliang Yu
摘要
Recent diffusion-based generative models achieve remarkable results by training on massive datasets, yet this practice raises concerns about memorization and copyright infringement. A proposed remedy is to train exclusively on noisy data with potential copyright issues, ensuring the model never observes original content. However, through the lens of deconvolution theory, we show that although it is theoretically feasible to learn the data distribution from noisy samples, the practical challenge of collecting sufficient samples makes successful learning nearly unattainable. To overcome this limitation, we propose to pretrain the model with a small fraction of clean data to guide the deconvolution process. Combined with our Stochastic Forward-Backward Deconvolution (SFBD) method, we attain FID 6.31 on CIFAR-10 with just 4% clean images (and 3.58 with 10%). We theoretically show that SFBD guides the model to learn the true data distribution. The result also highlights the importance of pretraining on limited but clean data or the alternative from similar datasets. Empirical studies further support these findings and offer additional insights.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Generative Modeling from Black-Box Corruptions via Self-Consistent Stochastic InterpolantsChirag Modi, Jiequn Han, Eric Vanden-Eijnden, Joan BrunaICLR 2026 · 被引用 4 次
- Ambient Dataloops: Generative Models for Dataset RefinementAdrian Rodriguez-Munoz, William Daspit, Adam Klivans, Antonio Torralba 等ICML 2026
- SFBD-OMNI: Bridge models for lossy measurement restoration with limited clean samplesHaoye Lu, Yaoliang Yu, Darren LoICLR 2026
- MAD: Manifold Attracted DiffusionDennis Elbrächter, Giovanni S. Alberti, Matteo SantacesariaICML 2026
它引用的顶会 Paper22
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
相关 Paper
- Does Generation Require Memorization? Creative Diffusion Models using Ambient DiffusionKulin Shah, Alkis Kalavasis, Adam R. Klivans, Giannis DarasICML 2025
- Ambient Diffusion: Learning Clean Distributions from Corrupted DataGiannis Daras, Kulin Shah, Yuval Dagan, Aravind Gollakota 等NeurIPS 2023 · 被引用 141 次
- Forward-Learned Discrete Diffusion: Learning how to noise to denoise fasterGrigory Bartosh, Teodora Pandeva, Sushrut Karmalkar, Javier ZazoICLR 2026 · 被引用 4 次
- Truncated Diffusion Probabilistic Models and Diffusion-based Adversarial Auto-EncodersHuangjie Zheng, Pengcheng He, Weizhu Chen, Mingyuan ZhouICLR 2023 · 被引用 19 次
- Self-diffusion for Solving Inverse ProblemsGuanxiong Luo, Shoujin HuangNeurIPS 2025 · 被引用 5 次
