Diffusion bridges vector quantized variational autoencoders
Max Cohen, Guillaume Quispe, Sylvain Le Corff, Charles Ollion, Eric Moulines
Abstract
Vector Quantized-Variational AutoEncoders (VQ-VAE) are generative models based on discrete latent representations of the data, where inputs are mapped to a finite set of learned embeddings. To generate new samples, an autoregressive prior distribution over the discrete states must be trained separately. This prior is generally very complex and leads to slow generation. In this work, we propose a new model to train the prior and the encoder/decoder networks simultaneously. We build a diffusion bridge between a continuous coded vector and a non-informative prior distribution. The latent discrete states are then given as random functions of these continuous vectors. We show that our model is competitive with the autoregressive prior on the mini-Imagenet and CIFAR dataset and is efficient in both optimization and sampling. Our framework also extends the standard VQ-VAE and enables end-to-end training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 252039c7-cf20-49f6-9014-bb50b79d56feCited by top-tier papers6
- A Continuous Time Framework for Discrete Denoising ModelsAndrew Campbell, Joe Benton, Valentin De Bortoli, Thomas Rainforth et al.NeurIPS 2022 · 496 citations
- Stochastic Segmentation with Conditional Categorical Diffusion ModelsLukas Zbinden, Lars Doorenbos, Theodoros Pissas, Adrian Thomas Huber et al.ICCV 2023 · 57 citations
- Formulating Discrete Probability Flow Through Optimal TransportPengze Zhang, Hubery Yin, Chen Li, Xiaohua XieNeurIPS 2023 · 11 citations
- Score-based Continuous-time Discrete Diffusion ModelsHaoran Sun, Lijun Yu, Bo Dai, Dale Schuurmans et al.ICLR 2023 · 7 citations
- SketchDNN: Joint Continuous-Discrete Diffusion for CAD Sketch GenerationSathvik Chereddy, John FemianiICML 2025
Builds on6
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
- Score-based Generative Modeling in Latent SpaceArash Vahdat, Karsten Kreis, Jan KautzNeurIPS 2021 · 903 citations
- Argmax Flows and Multinomial Diffusion: Learning Categorical DistributionsEmiel Hoogeboom, Didrik Nielsen, Priyank Jaini, Patrick Forré et al.NeurIPS 2021 · 782 citations
- Vector Quantized Diffusion Model for Text-to-Image SynthesisShuyang Gu, Dong Chen, Jianmin Bao, Fang Wen et al.CVPR 2022 · 607 citations
Related papers
- Global Context with Discrete Diffusion in Vector Quantised Modelling for Image GenerationMinghui Hu, Yujie Wang, Tat-Jen Cham, Jianfei Yang et al.CVPR 2022
- Autoregressive Image Generation without Vector QuantizationTianhong Li, Yonglong Tian, He Li, Mingyang Deng et al.NeurIPS 2024 · 758 citations
- Shape your Space: A Gaussian Mixture Regularization Approach to Deterministic AutoencodersAmrutha Saseendran, Kathrin Skubch, Stefan Falkner, Margret KeuperNeurIPS 2021 · 13 citations
- Restructuring Vector Quantization with the Rotation TrickChristopher Fifty, Ronald Guenther Junkins, Dennis Duan, Aniketh Iyengar et al.ICLR 2025 · 1 citation
- SQ-VAE: Variational Bayes on Discrete Representation with Self-annealed Stochastic QuantizationYuhta Takida, Takashi Shibuya, Wei-Hsiang Liao, Chieh-Hsin Lai et al.ICML 2022 · 99 citations
