Your Diffusion Model is Secretly a Noise Classifier and Benefits from Contrastive Training
Yunshu Wu, Yingtao Luo, Xianghao Kong, Vagelis Papalexakis, Greg Ver Steeg
Abstract
Diffusion models learn to denoise data and the trained denoiser is then used to generate new samples from the data distribution. In this paper, we revisit the diffusion sampling process and identify a fundamental cause of sample quality degradation: the denoiser is poorly estimated in regions that are far Outside Of the training Distribution (OOD), and the sampling process inevitably evaluates in these OOD regions. This can become problematic for all sampling methods, especially when we move to parallel sampling which requires us to initialize and update the entire sample trajectory of dynamics in parallel, leading to many OOD evaluations. To address this problem, we introduce a new self-supervised training objective that differentiates the levels of noise added to a sample, leading to improved OOD denoising performance. The approach is based on our observation that diffusion models implicitly define a log-likelihood ratio that distinguishes distributions with different amounts of noise, and this expression depends on denoiser performance outside the standard training distribution. We show by diverse experiments that the proposed contrastive diffusion training is effective for both sequential and parallel settings, and it improves the performance and speed of parallel samplers significantly.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80ec09ef-c968-4661-ae8f-e906dcd0ccfdCited by top-tier papers3
- Diffusion and Flow-based Copulas: Forgetting and Remembering DependenciesDavid Huk, Theodoros DamoulasICLR 2026 · 7 citations
- The Accumulation of Score Estimation Error in Diffusion ModelsBaoxiang He, Valentio Iverson, Shuai Li, Cheng Chen et al.ICML 2026
- Maximum Entropy Reinforcement Learning with Diffusion PolicyXiaoyi Dong, Jian Cheng, Xi Sheryl ZhangICML 2025
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
Related papers
- Consistent Diffusion Models: Mitigating Sampling Drift by Learning to be ConsistentGiannis Daras, Yuval Dagan, Alex Dimakis, Constantinos DaskalakisNeurIPS 2023 · 79 citations
- Interpreting and Improving Diffusion Models from an Optimization PerspectiveFrank Permenter, Chenyang YuanICML 2024 · 16 citations
- Temporal Difference Learning for Diffusion ModelsQizhen Ying, Yangchen Pan, Victor Prisacariu, Junfeng WenICML 2026
- A Continuous Time Framework for Discrete Denoising ModelsAndrew Campbell, Joe Benton, Valentin De Bortoli, Thomas Rainforth et al.NeurIPS 2022 · 496 citations
- Distributional Diffusion Models with Scoring RulesValentin De Bortoli, Alexandre Galashov, J. Swaroop Guntupalli, Guangyao Zhou et al.ICML 2025
