Generating images with sparse representations
Charlie Nash, Jacob Menick, Sander Dieleman, Peter W. Battaglia
Abstract
The high dimensionality of images presents architecture and sampling-efficiency challenges for likelihood-based generative models. Previous approaches such as VQ-VAE use deep autoencoders to obtain compact representations, which are more practical as inputs for likelihood-based models. We present an alternative approach, inspired by common image compression methods like JPEG, and convert images to quantized discrete cosine transform (DCT) blocks, which are represented sparsely as a sequence of DCT channel, spatial location, and DCT coefficient triples. We propose a Transformer-based autoregressive architecture, which is trained to sequentially predict the conditional distribution of the next element in such sequences, and which scales effectively to high resolution images. On a range of image datasets, we demonstrate that our approach can generate high quality, diverse images, with sample metric scores competitive with state of the art methods. We additionally show that simple modifications to our method yield effective image colorization and super-resolution models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dec4b313-cf87-4b7a-bd99-c22ce804cd14Cited by top-tier papers149
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Vector-quantized Image Modeling with Improved VQGANJiahui Yu, Xin Li, Jing Yu Koh, Han Zhang et al.ICLR 2022 · 753 citations
- Representation Alignment for Diffusion Transformers without External ComponentsDengyang Jiang, Mengmeng Wang, Liuzhuozheng Li, Lei Zhang et al.ICLR 2026 · 532 citations
- MaskGIT: Masked Generative Image TransformerHuiwen Chang, Han Zhang, Lu Jiang, Ce Liu et al.CVPR 2022 · 346 citations
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- Stabilizing Transformers for Reinforcement LearningEmilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu et al.ICML 2020 · 464 citations
Related papers
- Autoregressive Image Generation using Residual QuantizationDoyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho et al.CVPR 2022 · 184 citations
- Towards Accurate Image Coding: Improved Autoregressive Image Generation with Dynamic Vector QuantizationMengqi Huang, Zhendong Mao, Zhuowei Chen, Yongdong ZhangCVPR 2023
- Draft-and-Revise: Effective Image Generation with Contextual RQ-TransformerDoyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho et al.NeurIPS 2022 · 36 citations
- A Spark of Vision-Language Intelligence: 2-Dimensional Autoregressive Transformer for Efficient Finegrained Image GenerationLiang Chen, Sinan Tan, Zefan Cai, Weichu Xie et al.ICLR 2025
- Locally Hierarchical Auto-Regressive Modeling for Image GenerationTackgeun You, Saehoon Kim, Chiheon Kim, Doyup Lee et al.NeurIPS 2022 · 17 citations
