Randomized Autoregressive Visual Generation
Qihang Yu, Ju He, Xueqing Deng, Xiaohui Shen, Liang-Chieh Chen
Abstract
This paper presents Randomized AutoRegressive modeling (RAR) for visual generation, which sets a new state-of-theart performance on the image generation task while maintaining full compatibility with language modeling frameworks. The proposed RAR is simple: during a standard autoregressive training process with a next-token prediction objective, the input sequence-typically ordered in raster form-is randomly permuted into different factorization orders with a probability r, where r starts at 1 and linearly decays to 0 over the course of training. This annealing training strategy enables the model to learn to maximize the expected likelihood over all factorization orders and thus effectively improve the model's capability of modeling bidirectional contexts. Importantly, RAR preserves the integrity of the autoregressive modeling framework, ensuring full compatibility with language modeling while significantly improving performance in image generation. On the ImageNet-256 benchmark, RAR achieves an FID score of 1.48, not only surpassing prior state-of-the-art autoregressive image generators but also outperforming leading diffusion-based and masked transformer-based methods. Code and models will be made available at https: //github.com/bytedance/1d-tokenizer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a3b22191-8a18-405d-a0ee-e5cadd54aef6Cited by top-tier papers71
- Improved Mean Flows: On the Challenges of Fastforward Generative ModelsZhengyang Geng, Yiyang Lu, Zongze Wu, Eli Shechtman et al.CVPR 2026 · 116 citations
- DDT: Decoupled Diffusion TransformerShuai Wang, Zhi Tian, Weilin Huang, Limin WangCVPR 2026 · 102 citations
- Diffusion Beats Autoregressive in Data-Constrained SettingsMihir Prabhudesai, Mengning Wu, Amir Zadeh, Katerina Fragkiadaki et al.NeurIPS 2025 · 69 citations
- Puppeteer: Rig and Animate Your 3D ModelsChaoyue Song, Xiu Li, Fan Yang, Zhongcong Xu et al.NeurIPS 2025 · 48 citations
- GoT-R1: Unleashing Reasoning Capability of Autoregressive Visual Generation with Reinforcement LearningChengqi Duan, Rongyao Fang, Yuqing Wang, Kun Wang et al.ICLR 2026 · 43 citations
Builds on28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- reAR: Rethinking Visual Autoregressive Models via Token-wise Consistency RegularizationQiyuan He, Yicong Li, Haotian Ye, Jinghao Wang et al.ICLR 2026 · 4 citations
- Unified Multimodal Autoregressive Modeling with Shared Context—Visual Tokenizer is Key to UnificationWujian Peng, Lingchen Meng, Yuxuan Cai, Xianwei Zhuang et al.ICML 2026 · 2 citations
- D-AR: Diffusion via Autoregressive ModelsZiteng Gao, Mike Zheng ShouICLR 2026 · 11 citations
- End-to-End Autoregressive Image Generation with 1D Semantic TokenizerWenda Chu, Bingliang Zhang, Jiaqi Han, Yizhuo Li et al.ICML 2026 · 2 citations
- Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous TokensLijie Fan, Tianhong Li, Siyang Qin, Yuanzhen Li et al.ICLR 2025 · 1 citation
