HMAR: Efficient Hierarchical Masked Auto-Regressive Image Generation
Hermann Kumbong, Xian Liu, Tsung-Yi Lin, Ming-Yu Liu, Xihui Liu, Ziwei Liu, Daniel Y. Fu, Christopher Ré, David W. Romero
Abstract
Visual Auto-Regressive modeling (VAR) has shown promise in bridging the speed and quality gap between autoregressive image models and diffusion models. VAR reformulates autoregressive modeling by decomposing an image into successive resolution scales. During inference, an image is generated by predicting all the tokens in the next (higher-resolution) scale, conditioned on all tokens in all previous (lower-resolution) scales. However, this formulation suffers from reduced image quality due to the parallel generation of all tokens in a resolution scale; has sequence lengths scaling superlinearly in image resolution; and requires retraining to change the sampling schedule. We introduce Hierarchical Masked AutoRegressive modeling (HMAR), a new image generation algorithm that alleviates these issues using next-scale prediction and masked prediction to generate high-quality images with fast sampling. HMAR reformulates next-scale prediction as a Markovian process, wherein the prediction of each resolution scale is conditioned only on tokens in its immediate predecessor instead of the tokens in all predecessor resolutions. When predicting a resolution scale, HMAR uses a controllable multi-step masked generation procedure to generate a subset of the tokens in each step. On * Equal contribution. † Equal senior authorship. ImageNet 256×256 and 512×512 benchmarks, HMAR models match or outperform parameter-matched VAR, diffusion, and autoregressive baselines. We develop efficient IO-aware blocksparse attention kernels that allow HMAR to achieve faster training and inference times over VAR by over 2.5× and 1.75× respectively, as well as over 3× lower inference memory footprint. Finally, HMAR yields additional flexibility over VAR; its sampling schedule can be changed without further training, and it can be applied to image editing tasks in a zero-shot manner.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 562ad3eb-afa4-4c7a-a92a-5324e02fba49Cited by top-tier papers5
- Visual Autoregressive Modeling for Instruction-Guided Image EditingQingyang Mao, Qi Cai, Yehao Li, Yingwei Pan et al.ICLR 2026 · 21 citations
- Markovian Scale Prediction: A New Era of Visual Autoregressive GenerationYu Zhang, Jingyi Liu, Yiwei Shi, Qi Zhang et al.CVPR 2026 · 4 citations
- ClusterMark: Towards Robust Watermarking for Autoregressive Image Generators with Visual Token ClusteringDenis Lukovnikov, Andreas Müller, Erwin Quiring, Asja FischerCVPR 2026 · 3 citations
- SSG: Scaled Spatial Guidance for Multi-Scale Visual Autoregressive GenerationYoungwoo Shin, Jiwan Hur, Junmo KimICLR 2026 · 1 citation
- UniCompress: Token Compression for Unified Vision-Language Understanding and GenerationZiyao Wang, Chen Chen, Jingtao Li, Weiming Zhuang et al.CVPR 2026 · 1 citation
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- FastVAR: Linear Visual Autoregressive Modeling Via Cached Token PruningHang Guo, Yawei Li, Taolin Zhang, Jiangshan Wang et al.ICCV 2025 · 5 citations
- Frequency-Aware Autoregressive Modeling for Efficient High-Resolution Image SynthesisZhuokun Chen, Jugang Fan, Zhuowei Yu, Bohan Zhuang et al.ICCV 2025 · 3 citations
- SparVAR: Exploring Sparsity in Visual AutoRegressive Modeling for Training-Free AccelerationZekun Li, Ning Wang, Tongxin Bai, Changwang Mei et al.CVPR 2026 · 4 citations
- LazyVAR: Accelerating Visual Autoregressive Models via Scale-wise Token Pruning and Parallel Group DecodingRongge Mao, Chengqi Dong, S Kevin ZhouCVPR 2026
- Hierarchical Masked Autoregressive Models with Low-Resolution Token PivotsGuangting Zheng, Yehao Li, Yingwei Pan, Jiajun Deng et al.ICML 2025
