Autoregressive Image Generation with Randomized Parallel Decoding
Haopeng Li, Jinyue Yang, Guoqi Li, Huan Wang
摘要
We introduce ARPG, a novel visual Autoregressive model that enables Randomized Parallel Generation, addressing the inherent limitations of conventional raster-order approaches, which hinder inference efficiency and zero-shot generalization due to their sequential, predefined token generation order. Our key insight is that effective random-order modeling necessitates explicit guidance for determining the position of the next predicted token. To this end, we propose a novel decoupled decoding framework that decouples positional guidance from content representation, encoding them separately as queries and key-value pairs. By directly incorporating this guidance into the causal attention mechanism, our approach enables fully random-order training and generation, eliminating the need for bidirectional attention. Consequently, ARPG readily generalizes to zero-shot tasks such as image in-painting, out-painting, and resolution expansion. Furthermore, it supports parallel inference by concurrently processing multiple queries using a shared KV cache. On the ImageNet-1K 256 benchmark, our approach attains an FID of 1.83 with only 32 sampling steps, achieving over a 30 times speedup in inference and and a 75 percent reduction in memory consumption compared to representative recent autoregressive models at a similar scale.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Show-o2: Improved Native Unified Multimodal ModelsJinheng Xie, Zhenheng Yang, Mike Zheng ShouNeurIPS 2025 · 被引用 261 次
- OBS-Diff: Accurate Pruning For Diffusion Models in One-ShotJunhan Zhu, Hesong Wang, Mingluo Su, Zefang Wang 等ICLR 2026 · 被引用 26 次
- Locality-aware Parallel Decoding for Efficient Autoregressive Image GenerationZhuoyang Zhang, Luke J. Huang, Chengyue Wu, Shang Yang 等ICLR 2026 · 被引用 8 次
- Parallel Jacobi Decoding for Fast Autoregressive Image GenerationBoya Liao, Ying Li, Siyong Jian, Huan WangCVPR 2026 · 被引用 2 次
- SJD-SV: Speculative Jacobi Decoding with Semantics Verification for Autoregressive Image GenerationBaoquan Zhang, Bingqi Shan, Shihao Fang, Kenghong Lin 等ICML 2026
它引用的顶会 Paper21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- ZipAR: Parallel Autoregressive Image Generation through Spatial LocalityYefei He, Feng Chen, Yuanyu He, Shaoxuan He 等ICML 2025
- Parallelized Autoregressive Visual GenerationYuqing Wang, Shuhuai Ren, Zhijie Lin, Yujin Han 等CVPR 2025
- Holistic Tokenizer for Autoregressive Image GenerationAnlin Zheng, Haochen Wang, Yucheng Zhao, Weipeng Deng 等ICCV 2025 · 被引用 11 次
- Neighboring Autoregressive Modeling for Efficient Visual GenerationYefei He, Yuanyu He, Shaoxuan He, Feng Chen 等ICCV 2025 · 被引用 3 次
- RandAR: Decoder-only Autoregressive Visual Generation in Random OrdersZiqi Pang, Tianyuan Zhang, Fujun Luan, Yunze Man 等CVPR 2025
