N2M: Bridging Navigation and Manipulation by Learning Pose Preference from Rollout
Kaixin Chai, Hyunjun Lee, Joseph Lim
摘要
Determining where to execute the manipulation policy is a fundamental challenge in mobile manipulation. Most approaches have formulated this as a geometric search problem, prioritizing physical reachability. However, given the high sensitivity of modern learning-based manipulation policies, geometric criteria alone are insufficient. Optimal performance requires base positioning that is aware of the policy's preference. While recent works have attempted to address this, they remain limited in practicality due to reliance on pre-built scene reconstruction and slow inference. In this work, we introduce N2M that systematically reformulates the approach to base positioning problem, naturally overcoming limitations of previous methods. Our key insight is that policy preferences are inherent to the local scene structure and can be effectively learned from the policy rollouts. Technically, we propose a novel viewpoint augmentation strategy that enables the model to learn robust, viewpoint-invariant pose preferences with remarkable data efficiency. Extensive experiments demonstrate that N2M achieves state-of-the-art performance, outperforming both non-policy-aware baselines and recent policy-aware alternatives. Furthermore, we provide a comprehensive analysis highlighting N2M’s broad applicability, generalization capabilities, and data efficiency. Project website: https://clvrai.github.io/N2M/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point ModelingXumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang 等CVPR 2022 · 被引用 684 次
- Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and PlanningMoo Jin Kim, Yihuai Gao, Tsung-Yi Lin, Yen-Chen Lin 等ICLR 2026 · 被引用 325 次
- Find What You Want: Learning Demand-conditioned Object Attribute Space for Demand-driven NavigationHongcheng Wang, Andy Guan Hong Chen, Xiaoqi Li, Mingdong Wu 等NeurIPS 2023 · 被引用 33 次
- CoWs on Pasture: Baselines and Benchmarks for Language-Driven Zero-Shot Object NavigationSamir Yitzhak Gadre, Mitchell Wortsman, Gabriel Ilharco, Ludwig Schmidt 等CVPR 2023
- Data Scaling Laws in Imitation Learning for Robotic ManipulationFanqi Lin, Yingdong Hu, Pingyue Sheng, Chuan Wen 等ICLR 2025
相关 Paper
- Affordance-Guided Coarse-to-Fine Exploration for Base Placement in Open-Vocabulary Mobile ManipulationTzu-Jung Lin, Jia-Fong Yeh, Hung-Ting Su, Chung-Yi Lin 等AAAI 2026
- MoManipVLA: Transferring Vision-language-action Models for General Mobile ManipulationZhenyu Wu, Yuheng Zhou, Xiuwei Xu, Ziwei Wang 等CVPR 2025
- Active Perception Meets Rule-Guided RL: A Two-Phase Approach for Precise Object Navigation in Complex EnvironmentsLiang Qin, Min Wang, Peiwei Li, Wengang Zhou 等ICCV 2025 · 被引用 6 次
- FoAM: Foresight-Augmented Multi-Task Imitation Policy for Robotic ManipulationLitao Liu, Wentao Wang, Yifan Han, Zhuoli Xie 等AAAI 2026 · 被引用 4 次
- MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile ManipulationPingrui Zhang, Xianqiang Gao, Yuhan Wu, Kehui Liu 等ICCV 2025
