Learning with a Mole: Transferable latent spatial representations for navigation without reconstruction
Guillaume Bono, Leonid Antsfeld, Assem Sadek, Gianluca Monaci, Christian Wolf
Abstract
Agents navigating in 3D environments require some form of memory, which should hold a compact and actionable representation of the history of observations useful for decision taking and planning. In most end-to-end learning approaches the representation is latent and usually does not have a clearly defined interpretation, whereas classical robotics addresses this with scene reconstruction resulting in some form of map, usually estimated with geometry and sensor models and/or learning. In this work we propose to learn an actionable representation of the scene independently of the targeted downstream task and without explicitly optimizing reconstruction. The learned representation is optimized by a blind auxiliary agent trained to navigate with it on multiple short sub episodes branching out from a waypoint and, most importantly, without any direct visual observation. We argue and show that the blindness property is important and forces the (trained) latent representation to be the only means for planning. With probing experiments we show that the learned representation optimizes navigability and not reconstruction. On downstream tasks we show that it is robust to changes in distribution, in particular the sim2real gap, which we evaluate with a real physical robot in a real office building, significantly improving performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a08fd427-71b3-497b-82f6-042041b527cdCited by top-tier papers4
- Multi-Object Navigation with dynamically learned neural implicit representationsPierre Marza, Laëtitia Matignon, Olivier Simonin, Christian WolfICCV 2023 · 32 citations
- Kinaema: a recurrent sequence model for memory and pose in motionMert Bülent Sariyildiz, Philippe Weinzaepfel, Guillaume Bono, Gianluca Monaci et al.NeurIPS 2025 · 3 citations
- Learning to Navigate Efficiently and Precisely in Real EnvironmentsGuillaume Bono, Hervé Poirier, Leonid Antsfeld, Gianluca Monaci et al.CVPR 2024
- Reasoning in Visual Navigation of End-to-end Trained Agents: A Dynamical Systems ApproachSteeven Janny, Hervé Poirier, Leonid Antsfeld, Guillaume Bono et al.CVPR 2025
Builds on17
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- Offline Reinforcement Learning as One Big Sequence Modeling ProblemMichael Janner, Qiyang Li, Sergey LevineNeurIPS 2021 · 950 citations
- Object Goal Navigation using Goal-Oriented Semantic ExplorationDevendra Singh Chaplot, Dhiraj Gandhi, Abhinav Gupta, Ruslan SalakhutdinovNeurIPS 2020 · 857 citations
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion FramesErik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee et al.ICLR 2020 · 608 citations
Related papers
- Emergence of Maps in the Memories of Blind Navigation AgentsErik Wijmans, Manolis Savva, Irfan Essa, Stefan Lee et al.ICLR 2023 · 17 citations
- Active Neural MappingZike Yan, Haoxiang Yang, Hongbin ZhaICCV 2023 · 37 citations
- Goal-Aware Prediction: Learning to Model What MattersSuraj Nair, Silvio Savarese, Chelsea FinnICML 2020 · 71 citations
- No RL, No Simulation: Learning to Navigate without NavigatingMeera Hahn, Devendra Singh Chaplot, Shubham Tulsiani, Mustafa Mukadam et al.NeurIPS 2021 · 98 citations
- Unsupervised Learning of Visual 3D Keypoints for ControlBoyuan Chen, Pieter Abbeel, Deepak PathakICML 2021 · 46 citations
