MVAR: Visual Autoregressive Modeling with Scale and Spatial Markovian Conditioning
Jinhua Zhang, Wei Long, Minghao Han, Weiyi You, Shuhang Gu
摘要
Essential to visual generation is efficient modeling of visual data priors. Conventional next-token prediction methods define the process as learning the conditional probability distribution of successive tokens. Recently, next-scale prediction methods redefine the process to learn the distribution over multi-scale representations, significantly reducing generation latency. However, these methods condition each scale on all previous scales and require each token to consider all preceding tokens, exhibiting scale and spatial redundancy. To better model the distribution by mitigating redundancy, we propose Markovian Visual AutoRegressive modeling (MVAR), a novel autoregressive framework that introduces scale and spatial Markov assumptions to reduce the complexity of conditional probability modeling. Specifically, we introduce a scale-Markov trajectory that only takes as input the features of adjacent preceding scale for next-scale prediction, enabling the adoption of a parallel training strategy that significantly reduces GPU memory consumption. Furthermore, we propose spatial-Markov attention, which restricts the attention of each token to a localized neighborhood of size k at corresponding positions on adjacent scales, rather than attending to every token across these scales, for the pursuit of reduced modeling complexity. Building on these improvements, we reduce the computational complexity of attention calculation from O(N 2 ) to O(N k), enabling training with just eight NVIDIA RTX 4090 GPUs and eliminating the need for KV cache during inference. Extensive experiments on ImageNet demonstrate that MVAR achieves comparable or superior performance with both small model trained from scratch and large fine-tuned models, while reducing the average GPU memory footprint by 3.0×. (a) Vanilla next-scale prediction (b) Scale Markovian conditioning (c) Scale&spatial Markovian con. (MVAR)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Markovian Scale Prediction: A New Era of Visual Autoregressive GenerationYu Zhang, Jingyi Liu, Yiwei Shi, Qi Zhang 等CVPR 2026 · 被引用 4 次
- Differentiable Vector Quantization for Rate-Distortion Optimization of Generative Image CompressionShiyin Jiang, Wei Long, Minghao Han, Zhenghao Chen 等CVPR 2026 · 被引用 3 次
- Taming Sampling Perturbations with Variance Expansion Loss for Latent Diffusion ModelsQifan Li, Xingyu Zhou, Jinhua Zhang, Weiyi You 等CVPR 2026 · 被引用 1 次
- RADAR: VQ-VAE Decoder of VAR is a Good Student for Restoring Against Degradation by AccelerationZiyang Wang, Yue Zhang, Mingdao Wang, Yasen Zhang 等CVPR 2026
- LazyVAR: Accelerating Visual Autoregressive Models via Scale-wise Token Pruning and Parallel Group DecodingRongge Mao, Chengqi Dong, S Kevin ZhouCVPR 2026
它引用的顶会 Paper26
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
相关 Paper
- SparVAR: Exploring Sparsity in Visual AutoRegressive Modeling for Training-Free AccelerationZekun Li, Ning Wang, Tongxin Bai, Changwang Mei 等CVPR 2026 · 被引用 4 次
- FastVAR: Linear Visual Autoregressive Modeling Via Cached Token PruningHang Guo, Yawei Li, Taolin Zhang, Jiangshan Wang 等ICCV 2025 · 被引用 5 次
- Frequency-Aware Autoregressive Modeling for Efficient High-Resolution Image SynthesisZhuokun Chen, Jugang Fan, Zhuowei Yu, Bohan Zhuang 等ICCV 2025 · 被引用 3 次
- HMAR: Efficient Hierarchical Masked Auto-Regressive Image GenerationHermann Kumbong, Xian Liu, Tsung-Yi Lin, Ming-Yu Liu 等CVPR 2025
- FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive ModelsSenmao Li, Kai Wang, Salman Khan, Fahad Khan 等ICML 2026 · 被引用 2 次
