SpectralAR: Spectral Autoregressive Visual Generation
Yuanhui Huang, Weiliang Chen, Wenzhao Zheng, Yueqi Duan, Jie Zhou, Jiwen Lu
Abstract
Autoregressive visual generation has garnered increasing attention due to its scalability and compatibility with other modalities compared with diffusion models. Most existing methods construct visual sequences as spatial patches for autoregressive generation. However, image patches are inherently parallel, contradicting the causal nature of autoregressive modeling. To address this, we propose a Spectral AutoRegressive (SpectralAR) visual generation framework, which realizes causality for visual sequences from the spectral perspective. Specifically, we first transform an image into ordered spectral tokens with Nested Spectral Tokenization, representing lower to higher frequency components. We then perform autoregressive generation in a coarse-to-fine manner with the sequences of spectral tokens. By considering different levels of detail in images, our SpectralAR achieves both sequence causality and token efficiency without bells and whistles. We conduct extensive experiments on ImageNet-1K for image reconstruction and autoregressive generation, and SpectralAR achieves 3.02 gFID with only 64 tokens and 310M parameters. Project page: https://huang-yh.github.io/spectralar/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 403e3b31-5838-4429-8738-1e7f0aada7bfCited by top-tier papers3
- CaTok: Taming Mean Flows for One-Dimensional Causal Image TokenizationYitong Chen, Zuxuan Wu, Xipeng Qiu, Yu-Gang JiangCVPR 2026 · 6 citations
- CART: Compositional AutoRegressive Transformer for Image GenerationSiddharth Roheda, Rohit Chowdhury, Aniruddha Bala, Rohan JaiswalAAAI 2026 · 3 citations
- Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font GenerationHaonan Cai, Yuxuan Luo, Zhouhui LianCVPR 2026 · 1 citation
Builds on28
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
Related papers
- Autoregression with Self-Token PredictionDengsheng Chen, Yangming Shi, Enhua WuICML 2026
- "Principal Components" Enable a New Language of ImagesXin Wen, Bingchen Zhao, Ismail Elezi, Jiankang Deng et al.ICCV 2025 · 1 citation
- reAR: Rethinking Visual Autoregressive Models via Token-wise Consistency RegularizationQiyuan He, Yicong Li, Haotian Ye, Jinghao Wang et al.ICLR 2026 · 4 citations
- NFIG: Multi-Scale Autoregressive Image Generation via Frequency OrderingZhihao Huang, Xi Qiu, Yukuo Ma, Yifu Zhou et al.NeurIPS 2025 · 20 citations
- D-AR: Diffusion via Autoregressive ModelsZiteng Gao, Mike Zheng ShouICLR 2026 · 11 citations
