MambaOut: Do We Really Need Mamba for Vision?
Weihao Yu, Xinchao Wang
Abstract
In memory of Kobe Bryant "What can I say, Mamba out." -Kobe Bryant's NBA farewell speech in 2016. Linear Linear Conv σ SSM Linear σ Linear Linear Conv Linear σ Gated CNN block (e.g. Our MambaOut) Mamba block (e.g. Vision Mamba) (a) 0 5 10 15 20 25 MACs (G) 76 78 80 82 84 86 ImageNet Top-1 Accuracy (%) 20M 40M 80M Model Size MambaOut Vision Mamba VMamba PlainMamba Accuracy vs. MACs vs. Model Size (b) Figure 1: (a) Architecture of Gated CNN [18] and Mamba [25] blocks (omitting Normalization and shortcut). The Mamba block extends the Gated CNN with an additional state space model (SSM). As will be conceptually discussed in Section 3, SSM is not necessary for image classification on ImageNet [19, 66]. To empirically verify this claim, we stack Gated CNN blocks to build a series of models named MambaOut. (b) MambaOut outperforms visual Mamba models, e.g., Vision Mamhba [104], VMamba [50] and PlainMamba [88], on ImageNet image classification. Abstract -Mamba, an architecture with RNN-like token mixer of state space model (SSM), was recently introduced to address the quadratic complexity of the attention mechanism and subsequently applied to vision tasks 1 . Nevertheless, the performance of Mamba for vision is often underwhelming when compared with convolutional and attention-based models. In this paper, we delve into the essence of Mamba, and conceptually conclude that Mamba is ideally suited for tasks with long-sequence and autoregressive characteristics. For vision tasks, as image classification does not align with either characteristic, we hypothesize that Mamba is not necessary for this task; Detection and segmentation tasks are also not autoregressive, yet they adhere to the long-sequence characteristic, so we believe it is still worthwhile to explore Mamba's potential for these tasks. To empirically verify our hypotheses, we construct a series of models named MambaOut through stacking Mamba blocks while removing their core token mixer, SSM. Experimental results strongly support our hypotheses. Specifically, our MambaOut model surpasses all visual Mamba models on ImageNet image classification, indicating that Mamba is indeed unnecessary for this task. As for detection and segmentation, MambaOut cannot match the performance of state-of-the-art visual Mamba models, demonstrating the potential of Mamba for long-sequence visual tasks. 1 The vision tasks we discuss in this paper include image classification on ImageNet [19, 66] , object detection & instance segmentation on COCO [48] and semantic segmentation on ADE20K [103] . Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fa4df52a-ae07-499b-83b7-d278e2a54165Cited by top-tier papers43
- DyG-Mamba: Continuous State Space Modeling on Dynamic GraphsDongyuan Li, Shiyin Tan, Ying Zhang, Ming Jin et al.NeurIPS 2025 · 43 citations
- VMoBA: Mixture-of-Block Attention for Video Diffusion ModelsJianzong Wu, Liang Hou, Haotian Yang, Ye Tian et al.ICLR 2026 · 36 citations
- Learning Truncated Causal History Model for Video RestorationAmirhosein Ghasemabadi, Muhammad Kamran Janjua, Mohammad Salameh, Di NiuNeurIPS 2024 · 28 citations
- DAMamba: Vision State Space Model with Dynamic Adaptive ScanTanzhe Li, Caoshuo Li, Jiayi Lyu, Hongjuan Pei et al.NeurIPS 2025 · 24 citations
- VSSD: Vision Mamba With Non-Causal State Space DualityYuheng Shi, Mingjia Li, Minjing Dong, Chang XuICCV 2025 · 20 citations
Builds on43
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
Related papers
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- Achilles' Heel of Mamba: Essential difficulties of the Mamba architecture demonstrated by synthetic dataTianyi Chen, Pengxiao Lin, Zhiwei Wang, Zhi-Qin John XuNeurIPS 2025 · 4 citations
- Mamba-Adaptor: State Space Model Adaptor for Visual RecognitionFei Xie, Jiahao Nie, Yujin Tang, Wenkang Zhang et al.CVPR 2025
- Autoregressive Pretraining with Mamba in VisionSucheng Ren, Xianhang Li, Haoqin Tu, Feng Wang et al.ICLR 2025
- Demystify Mamba in Vision: A Linear Attention PerspectiveDongchen Han, Ziyi Wang, Zhuofan Xia, Yizeng Han et al.NeurIPS 2024 · 287 citations
