Efficient-VQGAN: Towards High-Resolution Image Generation with Efficient Vision Transformers
Shiyue Cao, Yueqin Yin, Lianghua Huang, Yu Liu, Xin Zhao, Deli Zhao, Kaiqi Huang
Abstract
Vector-quantized image modeling has shown great potential in synthesizing high-quality images. However, generating high-resolution images remains a challenging task due to the quadratic computational overhead of the self-attention process. In this study, we seek to explore a more efficient two-stage framework for high-resolution image generation with improvements in the following three aspects. (1) Based on the observation that the first quantization stage has solid local property, we employ a local attention-based quantization model instead of the global attention mechanism used in previous methods, leading to better efficiency and reconstruction quality. (2) We emphasize the importance of multi-grained feature interaction during image generation and introduce an efficient attention mechanism that combines global attention (long-range semantic consistency within the whole image) and local attention (fined-grained details). This approach results in faster generation speed, higher generation fidelity, and improved resolution. (3) We propose a new generation pipeline incorporating autoencoding training and autoregressive generation strategy, demonstrating a better paradigm for image synthesis. Extensive experiments demonstrate the superiority of our approach in high-quality and high-resolution image reconstruction and generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d1c47a5-474f-4ff8-804a-76967332c69cCited by top-tier papers15
- An Image is Worth 32 Tokens for Reconstruction and GenerationQihang Yu, Mark Weber, Xueqing Deng, Xiaohui Shen et al.NeurIPS 2024 · 331 citations
- Diffusion Transformers with Representation AutoencodersBoyang Zheng, Nanye Ma, Shengbang Tong, Saining XieICLR 2026 · 288 citations
- StrokeNUWA - Tokenizing Strokes for Vector Graphic SynthesisZecheng Tang, Chenfei Wu, Zekai Zhang, Minheng Ni et al.ICML 2024 · 27 citations
- CAT: Content-Adaptive Image TokenizationJunhong Shen, Kushal Tirumala, Michihiro Yasunaga, Ishan Misra et al.NeurIPS 2025 · 17 citations
- PanoLlama: Generating Endless and Coherent Panoramas with Next-Token-Prediction LLMsTeng Zhou, Xiaoyu Zhang, Yongchuan TangICCV 2025 · 1 citation
Builds on26
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Autoregressive Image Generation using Residual QuantizationDoyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho et al.CVPR 2022 · 184 citations
- Towards Accurate Image Coding: Improved Autoregressive Image Generation with Dynamic Vector QuantizationMengqi Huang, Zhendong Mao, Zhuowei Chen, Yongdong ZhangCVPR 2023
- Generating Diverse Structure for Image Inpainting With Hierarchical VQ-VAEJialun Peng, Dong Liu, Songcen Xu, Houqiang LiCVPR 2021
- Not All Image Regions Matter: Masked Vector Quantization for Autoregressive Image GenerationMengqi Huang, Zhendong Mao, Quan Wang, Yongdong ZhangCVPR 2023
- Locally Hierarchical Auto-Regressive Modeling for Image GenerationTackgeun You, Saehoon Kim, Chiheon Kim, Doyup Lee et al.NeurIPS 2022 · 17 citations
