N2M: Bridging Navigation and Manipulation by Learning Pose Preference from Rollout
Kaixin Chai, Hyunjun Lee, Joseph Lim
Abstract
Determining where to execute the manipulation policy is a fundamental challenge in mobile manipulation. Most approaches have formulated this as a geometric search problem, prioritizing physical reachability. However, given the high sensitivity of modern learning-based manipulation policies, geometric criteria alone are insufficient. Optimal performance requires base positioning that is aware of the policy's preference. While recent works have attempted to address this, they remain limited in practicality due to reliance on pre-built scene reconstruction and slow inference. In this work, we introduce N2M that systematically reformulates the approach to base positioning problem, naturally overcoming limitations of previous methods. Our key insight is that policy preferences are inherent to the local scene structure and can be effectively learned from the policy rollouts. Technically, we propose a novel viewpoint augmentation strategy that enables the model to learn robust, viewpoint-invariant pose preferences with remarkable data efficiency. Extensive experiments demonstrate that N2M achieves state-of-the-art performance, outperforming both non-policy-aware baselines and recent policy-aware alternatives. Furthermore, we provide a comprehensive analysis highlighting N2M’s broad applicability, generalization capabilities, and data efficiency. Project website: https://clvrai.github.io/N2M/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e4c4aace-dd53-4721-bb2c-e223fdf075deBuilds on7
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point ModelingXumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang et al.CVPR 2022 · 684 citations
- Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and PlanningMoo Jin Kim, Yihuai Gao, Tsung-Yi Lin, Yen-Chen Lin et al.ICLR 2026 · 325 citations
- Find What You Want: Learning Demand-conditioned Object Attribute Space for Demand-driven NavigationHongcheng Wang, Andy Guan Hong Chen, Xiaoqi Li, Mingdong Wu et al.NeurIPS 2023 · 33 citations
- CoWs on Pasture: Baselines and Benchmarks for Language-Driven Zero-Shot Object NavigationSamir Yitzhak Gadre, Mitchell Wortsman, Gabriel Ilharco, Ludwig Schmidt et al.CVPR 2023
- Data Scaling Laws in Imitation Learning for Robotic ManipulationFanqi Lin, Yingdong Hu, Pingyue Sheng, Chuan Wen et al.ICLR 2025
Related papers
- Affordance-Guided Coarse-to-Fine Exploration for Base Placement in Open-Vocabulary Mobile ManipulationTzu-Jung Lin, Jia-Fong Yeh, Hung-Ting Su, Chung-Yi Lin et al.AAAI 2026
- MoManipVLA: Transferring Vision-language-action Models for General Mobile ManipulationZhenyu Wu, Yuheng Zhou, Xiuwei Xu, Ziwei Wang et al.CVPR 2025
- Active Perception Meets Rule-Guided RL: A Two-Phase Approach for Precise Object Navigation in Complex EnvironmentsLiang Qin, Min Wang, Peiwei Li, Wengang Zhou et al.ICCV 2025 · 6 citations
- FoAM: Foresight-Augmented Multi-Task Imitation Policy for Robotic ManipulationLitao Liu, Wentao Wang, Yifan Han, Zhuoli Xie et al.AAAI 2026 · 4 citations
- MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile ManipulationPingrui Zhang, Xianqiang Gao, Yuhan Wu, Kehui Liu et al.ICCV 2025
