MoManipVLA: Transferring Vision-language-action Models for General Mobile Manipulation
Zhenyu Wu, Yuheng Zhou, Xiuwei Xu, Ziwei Wang, Haibin Yan
Abstract
we utilize pre-trained VLA models to generate waypoints of the end-effector with high generalization ability. We design motion planning objectives for the mobile base and the robot arm, which aim at maximizing the physical feasibility of the trajectory. Finally, we present an efficient bilevel objective optimization framework for trajectory generation, where the upper-level optimization predicts waypoints for base movement to enhance the manipulator policy space, and the lower-level optimization selects the optimal end-effector trajectory to complete the manipulation task. Extensive experimental results on OVMM and the real world demonstrate that MoManipVLA achieves a 4.2% higher success rate than the state-of-the-art mobile manipulation, and only requires 50 training cost for real world deployment due to the strong generalization ability in the pre-trained VLA models. Our project page can be found here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bcac245b-2dd5-4def-aa58-5d6279991d6cCited by top-tier papers7
- MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied NavigationXun Huang, Shijia Zhao, Yunxiang Wang, Xin Lu et al.CVPR 2026 · 19 citations
- RoboInter: A Holistic Intermediate Representation Suite Towards Robotic ManipulationHao Li, Ziqin Wang, Zi-han Ding, Shuai Yang et al.ICLR 2026 · 17 citations
- N2M: Bridging Navigation and Manipulation by Learning Pose Preference from RolloutKaixin Chai, Hyunjun Lee, Joseph LimICML 2026 · 3 citations
- Language-Grounded Decoupled Action Representation for Robotic ManipulationWuDing Weng, Tongshu Wu, Liucheng Chen, Siyu xie et al.CVPR 2026 · 2 citations
- From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot ManipulationYajie Li, Bozhou Zhang, Chun Gu, Zipei Ma et al.ICML 2026 · 2 citations
Builds on9
- Vision-Language Foundation Models as Effective Robot ImitatorsXinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu et al.ICLR 2024 · 375 citations
- An Embodied Generalist Agent in 3D WorldJiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu et al.ICML 2024 · 361 citations
- SG-Nav: Online 3D Scene Graph Prompting for LLM-based Zero-shot Object NavigationHang Yin, Xiuwei Xu, Zhenyu Wu, Jie Zhou et al.NeurIPS 2024 · 215 citations
- Skill Transformer: A Monolithic Policy for Mobile ManipulationXiaoyu Huang, Dhruv Batra, Akshara Rai, Andrew SzotICCV 2023 · 34 citations
- Multi-skill Mobile Manipulation for Object RearrangementJiayuan Gu, Devendra Singh Chaplot, Hao Su, Jitendra MalikICLR 2023 · 10 citations
Related papers
- MoMaGen: Generating Demonstrations under Soft and Hard Constraints for Multi-Step Bimanual Mobile ManipulationChengshu Li, Mengdi Xu, Arpit Bahety, Hang Yin et al.ICLR 2026 · 17 citations
- Joint Navigation and Manipulation Planning with 3D Interaction ChainsKeming Zhang, Sixian Zhang, Xinhang Song, Hongyu Wang et al.ICML 2026
- Affordance-Guided Coarse-to-Fine Exploration for Base Placement in Open-Vocabulary Mobile ManipulationTzu-Jung Lin, Jia-Fong Yeh, Hung-Ting Su, Chung-Yi Lin et al.AAAI 2026
- OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data SynthesisJunting Chen, Haotian Liang, Lingxiao Du, Weiyun Wang et al.NeurIPS 2025 · 7 citations
- Active Perception Meets Rule-Guided RL: A Two-Phase Approach for Precise Object Navigation in Complex EnvironmentsLiang Qin, Min Wang, Peiwei Li, Wengang Zhou et al.ICCV 2025 · 6 citations
