DisCo-Diff: Enhancing Continuous Diffusion Models with Discrete Latents
Yilun Xu, Gabriele Corso, Tommi S. Jaakkola, Arash Vahdat, Karsten Kreis
Abstract
Diffusion models (DMs) have revolutionized generative learning. They utilize a diffusion process to encode data into a simple Gaussian distribution. However, encoding a complex, potentially multimodal data distribution into a single continuous Gaussian distribution arguably represents an unnecessarily challenging learning problem. We propose Discrete-Continuous Latent Variable Diffusion Models (DisCo-Diff) to simplify this task by introducing complementary discrete latent variables. We augment DMs with learnable discrete latents, inferred with an encoder, and train DM and encoder end-to-end. DisCo-Diff does not rely on pre-trained networks, making the framework universally applicable. The discrete latents significantly simplify learning the DM's complex noise-to-data mapping by reducing the curvature of the DM's generative ODE. An additional autoregressive transformer models the distribution of the discrete latents, a simple step because DisCo-Diff requires only few discrete variables with small codebooks. We validate DisCo-Diff on toy data, several image synthesis tasks as well as molecular docking, and find that introducing discrete latents consistently improves model performance. For example, DisCo-Diff achieves state-of-the-art FID scores on class-conditioned ImageNet-64/128 datasets with ODE sampler.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8dbff83e-2ded-44d2-9182-5ee02ea4eed1Cited by top-tier papers13
- Align Your Flow: Scaling Continuous-Time Flow Map DistillationAmirmojtaba Sabour, Sanja Fidler, Karsten KreisNeurIPS 2025 · 91 citations
- Rényi Diffusion ModelsYirong Shen, Lu GAN, Cong LingICML 2026 · 4 citations
- Compositional Discrete Latent Code for High Fidelity, Productive Diffusion ModelsSamuel Lavoie, Michael Noukhovitch, Aaron C. CourvilleNeurIPS 2025 · 3 citations
- Latent Stochastic InterpolantsSaurabh Singh, Dmitry LagunICLR 2026 · 2 citations
- Flow to the Mode: Mode-Seeking Diffusion Autoencoders for State-of-the-Art Image TokenizationKyle Sargent, Kyle Hsu, Justin Johnson, Li Fei-Fei et al.ICCV 2025 · 1 citation
Builds on43
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Latent Diffusion for Language GenerationJustin Lovelace, Varsha Kishore, Chao Wan, Eliot Shekhtman et al.NeurIPS 2023 · 177 citations
- Discrete Predictor-Corrector Diffusion Models for Image SynthesisJosé Lezama, Tim Salimans, Lu Jiang, Huiwen Chang et al.ICLR 2023
- Discrete Modeling via Boundary Conditional Diffusion ProcessesYuxuan Gu, Xiaocheng Feng, Lei Huang, Yingsheng Wu et al.NeurIPS 2024
- LDMol: A Text-to-Molecule Diffusion Model with Structurally Informative Latent Space Surpasses AR ModelsJinho Chang, Jong Chul YeICML 2025
- InfoDiffusion: Representation Learning Using Information Maximizing Diffusion ModelsYingheng Wang, Yair Schiff, Aaron Gokaslan, Weishen Pan et al.ICML 2023 · 64 citations
