Autoregressive Image Generation with Randomized Parallel Decoding
Haopeng Li, Jinyue Yang, Guoqi Li, Huan Wang
Abstract
We introduce ARPG, a novel visual Autoregressive model that enables Randomized Parallel Generation, addressing the inherent limitations of conventional raster-order approaches, which hinder inference efficiency and zero-shot generalization due to their sequential, predefined token generation order. Our key insight is that effective random-order modeling necessitates explicit guidance for determining the position of the next predicted token. To this end, we propose a novel decoupled decoding framework that decouples positional guidance from content representation, encoding them separately as queries and key-value pairs. By directly incorporating this guidance into the causal attention mechanism, our approach enables fully random-order training and generation, eliminating the need for bidirectional attention. Consequently, ARPG readily generalizes to zero-shot tasks such as image in-painting, out-painting, and resolution expansion. Furthermore, it supports parallel inference by concurrently processing multiple queries using a shared KV cache. On the ImageNet-1K 256 benchmark, our approach attains an FID of 1.83 with only 32 sampling steps, achieving over a 30 times speedup in inference and and a 75 percent reduction in memory consumption compared to representative recent autoregressive models at a similar scale.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 88924d2a-0c49-4779-8c4b-9bb4224b9e2eCited by top-tier papers5
- Show-o2: Improved Native Unified Multimodal ModelsJinheng Xie, Zhenheng Yang, Mike Zheng ShouNeurIPS 2025 · 261 citations
- OBS-Diff: Accurate Pruning For Diffusion Models in One-ShotJunhan Zhu, Hesong Wang, Mingluo Su, Zefang Wang et al.ICLR 2026 · 26 citations
- Locality-aware Parallel Decoding for Efficient Autoregressive Image GenerationZhuoyang Zhang, Luke J. Huang, Chengyue Wu, Shang Yang et al.ICLR 2026 · 8 citations
- Parallel Jacobi Decoding for Fast Autoregressive Image GenerationBoya Liao, Ying Li, Siyong Jian, Huan WangCVPR 2026 · 2 citations
- SJD-SV: Speculative Jacobi Decoding with Semantics Verification for Autoregressive Image GenerationBaoquan Zhang, Bingqi Shan, Shihao Fang, Kenghong Lin et al.ICML 2026
Builds on21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- ZipAR: Parallel Autoregressive Image Generation through Spatial LocalityYefei He, Feng Chen, Yuanyu He, Shaoxuan He et al.ICML 2025
- Parallelized Autoregressive Visual GenerationYuqing Wang, Shuhuai Ren, Zhijie Lin, Yujin Han et al.CVPR 2025
- Holistic Tokenizer for Autoregressive Image GenerationAnlin Zheng, Haochen Wang, Yucheng Zhao, Weipeng Deng et al.ICCV 2025 · 11 citations
- Neighboring Autoregressive Modeling for Efficient Visual GenerationYefei He, Yuanyu He, Shaoxuan He, Feng Chen et al.ICCV 2025 · 3 citations
- RandAR: Decoder-only Autoregressive Visual Generation in Random OrdersZiqi Pang, Tianyuan Zhang, Fujun Luan, Yunze Man et al.CVPR 2025
