Ambient Diffusion: Learning Clean Distributions from Corrupted Data
Giannis Daras, Kulin Shah, Yuval Dagan, Aravind Gollakota, Alex Dimakis, Adam R. Klivans
摘要
We present the first diffusion-based framework that can learn an unknown distribution using only highly-corrupted samples. This problem arises in scientific applications where access to uncorrupted samples is impossible or expensive to acquire. Another benefit of our approach is the ability to train generative models that are less likely to memorize individual training samples since they never observe clean training data. Our main idea is to introduce additional measurement distortion during the diffusion process and require the model to predict the original corrupted image from the further corrupted image. We prove that our method leads to models that learn the conditional expectation of the full uncorrupted image given this additional measurement corruption. This holds for any corruption process that satisfies some technical conditions (and in particular includes inpainting and compressed sensing). We train models on standard benchmarks (CelebA, CIFAR-10 and AFHQ) and show that we can learn the distribution even when all the training samples have of their pixels missing. We also show that we can finetune foundation models on small corrupted datasets (e.g. MRI scans with block corruptions) and learn the clean distribution without memorizing the training set.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper43
- Detecting, Explaining, and Mitigating Memorization in Diffusion ModelsYuxin Wen, Yuchen Liu, Chen Chen, Lingjuan LyuICLR 2024 · 被引用 103 次
- Learning Diffusion Priors from Observations by Expectation MaximizationFrançois Rozet, Gérôme Andry, François Lanusse, Gilles LouppeNeurIPS 2024 · 被引用 79 次
- Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMsAbhimanyu Hans, John Kirchenbauer, Yuxin Wen, Neel Jain 等NeurIPS 2024 · 被引用 65 次
- Schrodinger Bridge Flow for Unpaired Data TranslationValentin De Bortoli, Iryna Korshunova, Andriy Mnih, Arnaud DoucetNeurIPS 2024 · 被引用 55 次
- Consistent Diffusion Meets Tweedie: Training Exact Ambient Diffusion Models with Noisy DataGiannis Daras, Alex Dimakis, Constantinos DaskalakisICML 2024 · 被引用 45 次
它引用的顶会 Paper26
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- A Diffusion Model with State Estimation for Degradation-Blind Inverse ImagingLiya Ji, Zhefan Rao, Sinno Jialin Pan, Chenyang Lei 等AAAI 2024 · 被引用 5 次
- An Expectation-Maximization Algorithm for Training Clean Diffusion Models from Corrupted ObservationsWeimin Bai, Yifei Wang, Wenzheng Chen, He SunNeurIPS 2024 · 被引用 21 次
- Score Distillation Beyond Acceleration: Generative Modeling from Corrupted DataYasi Zhang, Tianyu Chen, Zhendong Wang, Ying Nian Wu 等ICLR 2026 · 被引用 2 次
- Normalization-equivariant Diffusion Models: Learning Posterior Samplers From Noisy And Partial MeasurementsBrett Levac, Jon Tamir, Marcelo Pereyra, Julián TachellaICML 2026 · 被引用 2 次
- Ambient Diffusion Posterior Sampling: Solving Inverse Problems with Diffusion Models Trained on Corrupted DataAsad Aali, Giannis Daras, Brett Levac, Sidharth Kumar 等ICLR 2025
