Markovian Scale Prediction: A New Era of Visual Autoregressive Generation
Yu Zhang, Jingyi Liu, Yiwei Shi, Qi Zhang, Duoqian Miao, Changwei Wang, Longbing Cao
Abstract
Visual AutoRegressive modeling (VAR) based on next-scale prediction has revitalized autoregressive visual generation. Although its full-context dependency, i.e., modeling all previous scales for next-scale prediction, facilitates more stable and comprehensive representation learning by leveraging complete information flow, the resulting computational inefficiency and substantial overhead severely hinder VAR's practicality and scalability. This motivates us to develop a new VAR model with better performance and efficiency without full-context dependency. To address this, we reformulate VAR as a non-full-context Markov process, proposing Markov-VAR. It is achieved via Markovian Scale Prediction: we treat each scale as a Markov state and introduce a sliding window that compresses certain previous scales into a compact history vector to compensate for historical information loss owing to non-full-context dependency. Integrating the history vector with the Markov state yields a representative dynamic state that evolves under a Markov process. Extensive experiments demonstrate that Markov-VAR is extremely simple yet highly effective: Compared to VAR on ImageNet, Markov-VAR reduces FID by 10.5% (256 256) and decreases peak memory consumption by 83.8% (1024 1024). We believe that Markov-VAR can serve as a foundation for future research on visual autoregressive generation and other downstream tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ea4d058b-f3ef-49c4-8870-604e06a80bc6Cited by top-tier papers2
- SIGMA: Selective-Interleaved Generation with Multi-Attribute TokensXiaoyan Zhang, Zechen Bai, Haofan Wang, Yiren SongCVPR 2026 · 3 citations
- Adaptive Visual Autoregressive Acceleration via Dual-Linkage Entropy AnalysisYu Zhang, Jingyi Liu, Feng Liu, Duoqian Miao et al.ICML 2026
Builds on37
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- MVAR: Visual Autoregressive Modeling with Scale and Spatial Markovian ConditioningJinhua Zhang, Wei Long, Minghao Han, Weiyi You et al.ICLR 2026 · 9 citations
- Visual Implicit Autoregressive ModelingPengfei Jiang, Jixiang Luo, Luxi Lin, Zhaohong Huang et al.ICML 2026
- Progressive Supernet Training for Efficient Visual Autoregressive ModelingXiaoyue Chen, Yuling Shi, Kaiyuan Li, Huandong Wang et al.CVPR 2026 · 7 citations
- FlowAR: Scale-wise Autoregressive Image Generation Meets Flow MatchingSucheng Ren, Qihang Yu, Ju He, Xiaohui Shen et al.ICML 2025
- FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive ModelsSenmao Li, Kai Wang, Salman Khan, Fahad Khan et al.ICML 2026 · 2 citations
