RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
Jiaming Liu, Mengzhen Liu, Zhenyu Wang, Pengju An, Xiaoqi Li, Kaichen Zhou, Senqiao Yang, Renrui Zhang, Yandong Guo, Shanghang Zhang
Abstract
A fundamental objective in robot manipulation is to enable models to comprehend visual scenes and execute actions. Although existing Vision-Language-Action (VLA) models for robots can handle a range of basic tasks, they still face challenges in two areas: (1) insufficient reasoning ability to tackle complex tasks, and (2) high computational costs for VLA model fine-tuning and inference. The recently proposed state space model (SSM) known as Mamba demonstrates promising capabilities in non-trivial sequence modeling with linear inference complexity. Inspired by this, we introduce RoboMamba, an end-to-end robotic VLA model that leverages Mamba to deliver both robotic reasoning and action capabilities, while maintaining efficient fine-tuning and inference. Specifically, we first integrate the vision encoder with Mamba, aligning visual tokens with language embedding through co-training, empowering our model with visual common sense and robotic-related reasoning. To further equip RoboMamba with SE(3) pose prediction abilities, we explore an efficient fine-tuning strategy with a simple policy head. We find that once RoboMamba possesses sufficient reasoning capability, it can acquire manipulation skills with minimal fine-tuning parameters (0.1% of the model) and time. In experiments, RoboMamba demonstrates outstanding reasoning capabilities on general and robotic evaluation benchmarks. Meanwhile, our model showcases impressive pose prediction results in both simulation and real-world experiments, achieving inference speeds 3 times faster than existing VLA models. Our project web page: https://sites.google.com/view/robomamba-web
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers34
- ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich ManipulationJiawen Yu, Hairuo Liu, Qiaojun Yu, Jieji Ren et al.NeurIPS 2025 · 150 citations
- EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action ModelsYantai Yang, Yuhao Wang, Zichen Wen, Luo Zhongwei et al.NeurIPS 2025 · 94 citations
- CogVLA: Cognition-Aligned Vision-Language-Action Models via Instruction-Driven Routing & SparsificationWei Li, Renshan Zhang, Rui Shao, Jie He et al.NeurIPS 2025 · 87 citations
- Fast-in-Slow: A Dual-System VLA Model Unifying Fast Manipulation within Slow ReasoningHao Chen, Jiaming Liu, Chenyang Gu, Zhuoyang Liu et al.NeurIPS 2025 · 74 citations
- MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot ManipulationRongyu Zhang, Menghang Dong, Yuan Zhang, Liang Heng et al.AAAI 2026 · 56 citations
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- Shaking Up VLMs: Comparing Transformers and Structured State Space Models for Vision & Language ModelingGeorgios Pantazopoulos, Malvina Nikandrou, Alessandro Suglia, Oliver Lemon et al.EMNLP 2024 · 3 citations
- Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient InferenceHan Zhao, Min Zhang, Wei Zhao, Pengxiang Ding et al.AAAI 2025 · 125 citations
- MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language TrackingXinqi Liu, Li Zhou, Zikun Zhou, Jianqiu Chen et al.CVPR 2025
- Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic ScenesZhiYuan Feng, Zhaolu Kang, Qijie Wang, Zhiying Du et al.ICLR 2026 · 23 citations
- GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous InstructionsHelong Huang, Min Cen, Kai Tan, Xingyue Quan et al.AAAI 2026 · 12 citations
