Global Context with Discrete Diffusion in Vector Quantised Modelling for Image Generation
Minghui Hu, Yujie Wang, Tat-Jen Cham, Jianfei Yang, Ponnuthurai N. Suganthan
Abstract
The integration of Vector Quantised Variational AutoEncoder (VQ-VAE) with autoregressive models as generation part has yielded high-quality results on image generation. However, the autoregressive models will strictly follow the progressive scanning order during the sampling phase. This leads the existing VQ series models to hardly escape the trap of lacking global information. Denoising Diffusion Probabilistic Models (DDPM) in the continuous domain have shown a capability to capture the global context, while generating high-quality images. In the discrete state space, some works have demonstrated the potential to perform text generation and low resolution image generation. We show that with the help of a content-rich discrete visual codebook from VQ-VAE, the discrete diffusion model can also generate high fidelity images with global context, which compensates for the deficiency of the classical autoregressive model along pixel space. Meanwhile, the integration of the discrete VAE with the diffusion model resolves the drawback of conventional autoregressive models being oversized, and the diffusion model which demands excessive time in the sampling process when generating images. It is found that the quality of the generated images is heavily dependent on the discrete visual codebook. Extensive experiments demonstrate that the proposed Vector Quantised Discrete Diffusion Model (VQ-DDM) is able to achieve comparable performance to top-tier methods with low complexity. It also demonstrates outstanding advantages over other vectors quantised with autoregressive models in terms of image inpainting tasks without additional training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e28f959-ce20-4f41-80e8-3264290315a9Cited by top-tier papers24
- Blended Latent DiffusionOmri Avrahami, Ohad Fried, Dani LischinskiSIGGRAPH 2023 · 339 citations
- MoVQ: Modulating Quantized Vectors for High-Fidelity Image GenerationChuanxia Zheng, Tung-Long Vuong, Jianfei Cai, Dinh PhungNeurIPS 2022 · 156 citations
- Online Clustered CodebookChuanxia Zheng, Andrea VedaldiICCV 2023 · 67 citations
- Stochastic Segmentation with Conditional Categorical Diffusion ModelsLukas Zbinden, Lars Doorenbos, Theodoros Pissas, Adrian Thomas Huber et al.ICCV 2023 · 57 citations
- Fast Solvers for Discrete Diffusion Models: Theory and Applications of High-Order AlgorithmsYinuo Ren, Haoxuan Chen, Yuchen Zhu, Wei Guo et al.NeurIPS 2025 · 51 citations
Builds on14
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
Related papers
- Vector Quantized Diffusion Model for Text-to-Image SynthesisShuyang Gu, Dong Chen, Jianmin Bao, Fang Wen et al.CVPR 2022 · 607 citations
- Diffusion bridges vector quantized variational autoencodersMax Cohen, Guillaume Quispe, Sylvain Le Corff, Charles Ollion et al.ICML 2022 · 16 citations
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
- PQGAN: Product-Quantised Image Representation for High-Quality Image SynthesisDenis Zavadski, Nikita Philip Tatsch, Carsten RotherICLR 2026
- Variational Autoencoding Discrete Diffusion with Enhanced Dimensional Correlations ModelingTianyu Xie, Shuchen Xue, Zijin Feng, Tianyang Hu et al.ICLR 2026 · 14 citations
