A Complete Recipe for Diffusion Generative Models
Kushagra Pandey, Stephan Mandt
Abstract
Score-based Generative Models (SGMs) have demonstrated exceptional synthesis outcomes across various tasks. However, the current design landscape of the forward diffusion process remains largely untapped and often relies on physical heuristics or simplifying assumptions. Utilizing insights from the development of scalable Bayesian posterior samplers, we present a complete recipe for formulating forward processes in SGMs, ensuring convergence to the desired target distribution. Our approach reveals that several existing SGMs can be seen as specific manifestations of our framework. Building upon this method, we introduce Phase Space Langevin Diffusion (PSLD), which relies on score-based modeling within an augmented space enriched by auxiliary variables akin to physical phase space. Empirical results exhibit the superior sample quality and improved speed-quality trade-off of PSLD compared to various competing approaches on established image synthesis benchmarks. Remarkably, PSLD achieves sample quality akin to state-of-the-art SGMs (FID: 2.10 for unconditional CIFAR-10 generation). Lastly, we demonstrate the applicability of PSLD in conditional synthesis using pre-trained score networks, offering an appealing alternative as an SGM backbone for future advancements. Code and model checkpoints can be accessed at https://github.com/mandt-lab/PSLD . We include the proofs for Theorems 2.1 and 2.2 in Appendix A.1 for completeness. These results provide a general recipe for designing forward processes in SGMs. For the SGM to be a useful forward process, we need it to converge to a simple factorized distribution that serves as the initialization point of the backwards (generative) process. Consequently, we consider the following form of the stationary distribution p s (z): p s (z) = N (x; 0 dx , I dx )N (0 dm , M I dm ). (5) This form results from setting U (x) = x T x 2 in Eqn. 3. Therefore, for a positive semidefinite matrix D(z) and a skew-symmetric matrix Q(z), the most general class of forward processes which lead to an invariant distribution p s (z) can be specified by substituting the form of ∇H(z) (corresponding to p s (z) defined in Eqn. 5) in Eqn. 4. A similar characterization of forward processes has also been explored in a concurrent work by [27] in the context of likelihood estimation (see Section 5). Additional constraints on D and Q
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd1f6ef7-759f-4cf3-a002-736a52167a8fCited by top-tier papers10
- Precipitation Downscaling with Spatiotemporal Video DiffusionPrakhar Srivastava, Ruihan Yang, Gavin Kerrigan, Gideon Dresdner et al.NeurIPS 2024 · 27 citations
- Fast samplers for Inverse Problems in Iterative Refinement modelsKushagra Pandey, Ruihan Yang, Stephan MandtNeurIPS 2024 · 19 citations
- Diffusion Rejection SamplingByeonghu Na, Yeongmin Kim, Minsang Park, DongHyeok Shin et al.ICML 2024 · 11 citations
- Variational Control for Guidance in Diffusion ModelsKushagra Pandey, Farrin Marouf Sofian, Felix Draxler, Theofanis Karaletsos et al.ICML 2025
- Heavy-Tailed Diffusion ModelsKushagra Pandey, Jaideep Pathak, Yilun Xu, Stephan Mandt et al.ICLR 2025
Builds on31
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Score-Based Generative Modeling with Critically-Damped Langevin DiffusionTim Dockhorn, Arash Vahdat, Karsten KreisICLR 2022 · 276 citations
- Wasserstein Convergence of Critically Damped Langevin DiffusionsStanislas Strasman, Sobihan Surendran, Claire Boyer, Sylvain Le Corff et al.NeurIPS 2025 · 5 citations
- Preconditioned Langevin Dynamics with Score-based Generative Models for Infinite-Dimensional Linear Bayesian Inverse ProblemsLorenzo Baldassari, Josselin Garnier, Knut Solna, Maarten V. de HoopNeurIPS 2025 · 5 citations
- Wavelet Score-Based Generative ModelingFlorentin Guth, Simon Coste, Valentin De Bortoli, Stéphane MallatNeurIPS 2022 · 98 citations
- Denoising MCMC for Accelerating Diffusion-Based Generative ModelsBeomsu Kim, Jong Chul YeICML 2023 · 18 citations
