FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching
Sucheng Ren, Qihang Yu, Ju He, Xiaohui Shen, Alan L. Yuille, Liang-Chieh Chen
Abstract
Autoregressive (AR) modeling has achieved remarkable success in natural language processing by enabling models to generate text with coherence and contextual understanding through next token prediction. Recently, in image generation, VAR proposes scale-wise autoregressive modeling, which extends the next token prediction to the next scale prediction, preserving the 2D structure of images. However, VAR encounters two primary challenges: (1) its complex and rigid scale design limits generalization in next scale prediction, and (2) the generator's dependence on a discrete tokenizer with the same complex scale structure restricts modularity and flexibility in updating the tokenizer. To address these limitations, we introduce FlowAR, a general next scale prediction method featuring a streamlined scale design, where each subsequent scale is simply double the previous one. This eliminates the need for VAR's intricate multi-scale residual tokenizer and enables the use of any off-the-shelf Variational AutoEncoder (VAE). Our simplified design enhances generalization in next scale prediction and facilitates the integration of Flow Matching for high-quality image synthesis. We validate the effectiveness of FlowAR on the challenging ImageNet-256 benchmark, demonstrating superior generation performance compared to previous methods. Codes will be available at https: //github.com/OliverRensu/FlowAR .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers21
- SANA-Video: Efficient Video Generation with Block Linear Diffusion TransformerJunsong Chen, Yuyang Zhao, Jincheng Yu, Ruihang Chu et al.ICLR 2026 · 96 citations
- VITA: Vision-to-Action Flow Matching PolicyDechen Gao, BOQI ZHAO, Andrew Lee, Ian Chuang et al.ICLR 2026 · 27 citations
- Next Patch Prediction for AutoRegressive Visual GenerationYatian Pang, Peng Jin, Shuo Yang, Bin Zhu et al.AAAI 2026 · 24 citations
- FlexVAR: Flexible Visual Autoregressive Modeling without Residual PredictionSiyu Jiao, Gengwei Zhang, Yinlong Qian, Jiancheng Huang et al.NeurIPS 2025 · 23 citations
- Latent Denoising Makes Good TokenizersJiawei Yang, Tianhong Li, Lijie Fan, Yonglong Tian et al.ICLR 2026 · 17 citations
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Beyond Next-Token: Next-X Prediction for Autoregressive Visual GenerationSucheng Ren, Qihang Yu, Ju He, Xiaohui Shen et al.ICCV 2025 · 8 citations
- FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive ModelsSenmao Li, Kai Wang, Salman Khan, Fahad Khan et al.ICML 2026 · 2 citations
- Fundamental Limits of Visual Autoregressive Transformers: Universal Approximation AbilitiesYifang Chen, Xiaoyu Li, Yingyu Liang, Zhenmei Shi et al.ICML 2025
- LazyVAR: Accelerating Visual Autoregressive Models via Scale-wise Token Pruning and Parallel Group DecodingRongge Mao, Chengqi Dong, S Kevin ZhouCVPR 2026
- DVAR: Dynamic Visual Autoregressive Modeling for Image Super-ResolutionYu Zheng, Kai Zhang, Wei Zhu, Qingguo Liu et al.CVPR 2026
