Autoregressive Pretraining with Mamba in Vision
Sucheng Ren, Xianhang Li, Haoqin Tu, Feng Wang, Fangxun Shu, Lei Zhang, Jieru Mei, Linjie Yang, Peng Wang, Heng Wang, Alan L. Yuille, Cihang Xie
Abstract
The vision community has started to build with the recently developed state space model, Mamba, as the new backbone for a range of tasks. This paper shows that Mamba's visual capability can be significantly enhanced through autoregressive pretraining, a direction not previously explored. Efficiency-wise, the autoregressive nature can well capitalize on the Mamba's unidirectional recurrent structure, enabling faster overall training speed compared to other training strategies like mask modeling. Performance-wise, autoregressive pretraining equips the Mamba architecture with markedly higher accuracy over its supervised-trained counterparts and, more importantly, successfully unlocks its scaling potential to large and even huge model sizes. For example, with autoregressive pretraining, a base-size Mamba attains 83.2% ImageNet accuracy, outperforming its supervised counterpart by 2.0%; our huge-size Mamba, the largest Vision Mamba to date, attains 85.0% ImageNet accuracy (85.5% when finetuned with 384 × 384 inputs), notably surpassing all other Mamba variants in vision. The code is available at https://github.com/OliverRensu/ARM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 26e66fba-3ed7-4438-b2ad-d92918ec3efcCited by top-tier papers13
- RoMA: Scaling up Mamba-based Foundation Models for Remote SensingFengxiang Wang, Yulin Wang, Mingshuo Chen, Haotian Wang et al.NeurIPS 2025 · 14 citations
- REOrdering Patches Improves Vision ModelsDeclan Kutscher, David M. Chan, Yutong Bai, Trevor Darrell et al.NeurIPS 2025 · 3 citations
- State Space Prompting via Gathering and Spreading Spatio-Temporal Information for Video UnderstandingJiahuan Zhou, Kai Zhu, Zhenyu Cui, Zichen Liu et al.NeurIPS 2025 · 2 citations
- RNN as Linear Transformer: A Closer Investigation into Representational Potentials of Visual Mamba ModelsTiming Yang, Feng Wang, Guoyizhe WeiCVPR 2026 · 2 citations
- Semi-ViM: Bidirectional State Space Model for Mitigating Label Imbalance in Semi-Supervised LearningHongyang He, Hongyang Xie, Haochen You, Victor SanchezICCV 2025 · 1 citation
Builds on34
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
Related papers
- MAP: Unleashing Hybrid Mamba-Transformer Vision Backbone's Potential with Masked Autoregressive PretrainingYunze Liu, Li YiCVPR 2025
- MambaOut: Do We Really Need Mamba for Vision?Weihao Yu, Xinchao WangCVPR 2025
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- Mamba-Reg: Vision Mamba Also Needs RegistersFeng Wang, Jiahao Wang, Sucheng Ren, Guoyizhe Wei et al.CVPR 2025
- Scalable Pre-training of Large Autoregressive Image ModelsAlaaeldin El-Nouby, Michal Klein, Shuangfei Zhai, Miguel Ángel Bautista et al.ICML 2024 · 130 citations
