Spectral Image Tokenizer
Carlos Esteves, Mohammed Suhail, Ameesh Makadia
Abstract
Image tokenizers map images to sequences of discrete tokens, and are a crucial component of autoregressive transformer-based image generation. The tokens are typically associated with spatial locations in the input image, arranged in raster scan order, which is not ideal for autoregressive modeling. In this paper, we propose to tokenize the image spectrum instead, obtained from a discrete wavelet transform (DWT), such that the sequence of tokens represents the image in a coarse-to-fine fashion. Our tokenizer brings several advantages: 1) it leverages that natural images are more compressible at high frequencies, 2) it can take and reconstruct images of different resolutions without retraining, 3) it improves the conditioning for next-token prediction -- instead of conditioning on a partial line-by-line reconstruction of the image, it takes a coarse reconstruction of the full image, 4) it enables partial decoding where the first few generated tokens can reconstruct a coarse version of the image, 5) it enables autoregressive models to be used for image upsampling. We evaluate the tokenizer reconstruction metrics as well as multiscale image generation, text-guided image upsampling and editing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7b31b428-9c53-4aa9-9898-d23a200a7437Cited by top-tier papers8
- Latent Wavelet Diffusion For Ultra High-Resolution Image SynthesisLuigi Sigillo, Shengfeng He, Danilo ComminielloICLR 2026 · 8 citations
- Improving Progressive Generation with Decomposable Flow MatchingMoayed Haji-Ali, Willi Menapace, Ivan Skorokhodov, Arpit Sahni et al.NeurIPS 2025 · 7 citations
- SpectralAR: Spectral Autoregressive Visual GenerationYuanhui Huang, Weiliang Chen, Wenzhao Zheng, Yueqi Duan et al.ICCV 2025 · 2 citations
- Flow Along the K-Amplitude for Generative ModelingWeitao Du, Jiasheng Tang, Shuning Chang, Yu Rong et al.ICLR 2026 · 2 citations
- StrandDesigner: Towards Practical Strand Generation with Sketch GuidanceNa Zhang, Moran Li, Chengming Xu, Han Feng et al.ACM MM 2025 · 1 citation
Builds on21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng et al.NeurIPS 2024 · 1,199 citations
Related papers
- (1D) Ordered Tokens Enable Efficient Test-Time SearchZhitong Gao, Parham Rezaei, Ali Cy, Mingqiao Ye et al.ICML 2026 · 1 citation
- SSG: Scaled Spatial Guidance for Multi-Scale Visual Autoregressive GenerationYoungwoo Shin, Jiwan Hur, Junmo KimICLR 2026 · 1 citation
- Highly Compressed Tokenizer Can Generate Without TrainingLukas Lao Beyer, Tianhong Li, Xinlei Chen, Sertac Karaman et al.ICML 2025
- End-to-End Autoregressive Image Generation with 1D Semantic TokenizerWenda Chu, Bingliang Zhang, Jiaqi Han, Yizhuo Li et al.ICML 2026 · 2 citations
- ImageFolder: Autoregressive Image Generation with Folded TokensXiang Li, Kai Qiu, Hao Chen, Jason Kuen et al.ICLR 2025
