SPIN: Simultaneous Perception, Interaction and Navigation
Shagun Uppal, Ananye Agarwal, Haoyu Xiong, Kenneth Shaw, Deepak Pathak
摘要
While there has been remarkable progress recently in the fields of manipulation and locomotion, mobile manip-ulation remains a long-standing challenge. Compared to locomotion or static manipulation, a mobile system must make a diverse range of long-horizon tasks feasible in un-structured and dynamic environments. While the applications are broad and interesting, there are a plethora of chal-lenges in developing these systems such as coordination be-tween the base and arm, reliance on onboard perception for perceiving and interacting with the environment, and most importantly, simultaneously integrating all these parts to-gether. Prior works approach the problem using disentangled modular skills for mobility and manipulation that are trivially tied together. This causes several limitations such as compounding errors, delays in decision-making, and no whole-body coordination. In this work, we present a re-active mobile manipulation framework that uses an active visual system to consciously perceive and react to its en-vironment. Similar to how humans leverage whole-body and hand-eye coordination, we develop a mobile manipu-lator that exploits its ability to move and see, more specifically - to move in order to see and to see in order to move. This allows it to not only move around and interact with its environment but also, choose “when” to perceive “what” using an active visual system. We observe that such an agent learns to navigate around complex cluttered sce-narios while displaying agile whole-body coordination using only ego-vision without needing to create environment maps. Videos are available at https://spin-robot.github.io
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile ManipulationSixiang Chen, Jiaming Liu, Siyuan Qian, Han Jiang 等NeurIPS 2025 · 被引用 30 次
- Multi-Modal Grounded Planning and Efficient Replanning for Learning Embodied Agents with a Few ExamplesTaewoong Kim, Byeonghwi Kim, Jonghyun ChoiAAAI 2025 · 被引用 8 次
- Kinaema: a recurrent sequence model for memory and pose in motionMert Bülent Sariyildiz, Philippe Weinzaepfel, Guillaume Bono, Gianluca Monaci 等NeurIPS 2025 · 被引用 3 次
- APPLE: Toward General Active Perception via Reinforcement LearningTim Schneider, Cristiana de Farias, Roberto Calandra, Liming Chen 等ICLR 2026 · 被引用 2 次
- Better than Your Teacher: LLM Agents that learn from Privileged AI FeedbackSanjiban Choudhury, Paloma SodhiICLR 2025 · 被引用 1 次
它引用的顶会 Paper4
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion FramesErik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee 等ICLR 2020 · 被引用 608 次
- Learning To Explore Using Active Neural SLAMDevendra Singh Chaplot, Dhiraj Gandhi, Saurabh Gupta, Abhinav Gupta 等ICLR 2020 · 被引用 603 次
- RobustNav: Towards Benchmarking Robustness in Embodied NavigationPrithvijit Chattopadhyay, Judy Hoffman, Roozbeh Mottaghi, Aniruddha KembhaviICCV 2021 · 被引用 68 次
- Bayesian Imitation Learning for End-to-End Mobile ManipulationYuqing Du, Daniel Ho, Alex Alemi, Eric Jang 等ICML 2022 · 被引用 16 次
相关 Paper
- Active Vision Reinforcement Learning under Limited Visual ObservabilityJinghuan Shang, Michael S. RyooNeurIPS 2023 · 被引用 1 次
- ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon TasksKaijun Wang, Liqin Lu, Mingyu Liu, Jianuo Jiang 等AAAI 2026 · 被引用 6 次
- WholeBodyVLA: Towards Unified Latent VLA for Whole-body Loco-manipulation ControlHaoran Jiang, Jin Chen, Qingwen Bu, Li Chen 等ICLR 2026 · 被引用 56 次
- Uni-Sight: An E2E Vision-Language-Action System Unifying Multi-View Alignment and Multi-Modal FusionDaixun Li, Sibo He, Jiayun Tian, Yusi Zhang 等ACM MM 2025 · 被引用 1 次
- Multi-skill Mobile Manipulation for Object RearrangementJiayuan Gu, Devendra Singh Chaplot, Hao Su, Jitendra MalikICLR 2023 · 被引用 10 次
