Is Mapping Necessary for Realistic PointGoal Navigation?
Ruslan Partsey, Erik Wijmans, Naoki Yokoyama, Oles Dobosevych, Dhruv Batra, Oleksandr Maksymets
Abstract
Can an autonomous agent navigate in a new environment without building an explicit map? For the task of PointGoal navigation ('Go to Δx, Δy’) under idealized settings (no RGB-D and actuation noise, perfect GPS+Compass), the answer is a clear ‘yes' - mapless neural models composed of task-agnostic components (CNNs and RNNs) trained with large-scale reinforcement learning achieve 100% Success on a standard dataset (Gibson [24] ). However, for PointNav in a realistic setting (RGB-D and actuation noise, no GPS+Compass), this is an open question; one we tackle in this paper. The strongest published result for this task is 71.7% Success [39]. <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> According to Habitat Challenge 2020 PointNav benchmark held annually. A concurrent as-yet-unpublished result has reported 91% Success on 2021's benchmark, but we are unable to comment on the details because an associated report is not available.First, we identify the main (perhaps, only) cause of the drop in performance: absence of GPS+Compass. An agent with perfect GPS+Compass faced with RGB-D sensing and actuation noise achieves 99.8% Success (Gibson- v2 val). This suggests that (to paraphrase a meme) robust visual odometry is all we need for realistic PointNav; if we can achieve that, we can ignore the sensing and actuation noise. With that as our operating hypothesis, we scale dataset size, model size, and develop human-annotation-free dataaugmentation techniques to train neural models for visual odometry. We advance state of the art on the Habitat Realistic PointNav Challenge - SPL by 40% (relative), 53 to 74, and Success by 31% (relative), 71 to 94. While our approach does not saturate or ‘solve’ this dataset, this strong improvement combined with promising zero-shot sim2real transfer (to a LoCoBot robot) provides evidence consistent with the hypothesis that explicit mapping may not be necessary for navigation, even in a realistic setting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc0da49d-fee7-496d-9562-50f7ed397d22Cited by top-tier papers4
- Exploiting Proximity-Aware Tasks for Embodied Social NavigationEnrico Cancelli, Tommaso Campari, Luciano Serafini, Angel X. Chang et al.ICCV 2023 · 16 citations
- FloNa: Floor Plan Guided Embodied Visual NavigationJiaxin Li, Weiqi Huang, Zan Wang, Wei Liang et al.AAAI 2025 · 11 citations
- Towards Audio-Visual Navigation in Noisy Environments: A Large-Scale Benchmark Dataset and an Architecture Considering Multiple Sound-SourcesZhanbo Shi, Lin Zhang, Linfei Li, Ying ShenAAAI 2025 · 8 citations
- Phone2Proc: Bringing Robust Robots into Our Chaotic WorldMatt Deitke, Rose Hendrix, Ali Farhadi, Kiana Ehsani et al.CVPR 2023
Builds on6
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion FramesErik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee et al.ICLR 2020 · 608 citations
- Learning To Explore Using Active Neural SLAMDevendra Singh Chaplot, Dhiraj Gandhi, Saurabh Gupta, Abhinav Gupta et al.ICLR 2020 · 603 citations
- The Surprising Effectiveness of Visual Odometry Techniques for Embodied PointGoal NavigationXiaoming Zhao, Harsh Agrawal, Dhruv Batra, Alexander G. SchwingICCV 2021 · 50 citations
Related papers
- THDA: Treasure Hunt Data Augmentation for Semantic NavigationOleksandr Maksymets, Vincent Cartillier, Aaron Gokaslan, Erik Wijmans et al.ICCV 2021 · 107 citations
- Auxiliary Tasks and Exploration Enable ObjectGoal NavigationJoel Ye, Dhruv Batra, Abhishek Das, Erik WijmansICCV 2021 · 137 citations
- ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal EmbeddingsArjun Majumdar, Gunjan Aggarwal, Bhavika Devnani, Judy Hoffman et al.NeurIPS 2022 · 344 citations
- Emergence of Maps in the Memories of Blind Navigation AgentsErik Wijmans, Manolis Savva, Irfan Essa, Stefan Lee et al.ICLR 2023 · 17 citations
- No RL, No Simulation: Learning to Navigate without NavigatingMeera Hahn, Devendra Singh Chaplot, Shubham Tulsiani, Mustafa Mukadam et al.NeurIPS 2021 · 98 citations
