Reasoning in Visual Navigation of End-to-end Trained Agents: A Dynamical Systems Approach
Steeven Janny, Hervé Poirier, Leonid Antsfeld, Guillaume Bono, Gianluca Monaci, Boris Chidlovskii, Francesco Giuliari, Alessio Del Bue, Christian Wolf
Abstract
Progress in Embodied AI has made it possible for end-toend-trained agents to navigate in photo-realistic environments with high-level reasoning and zero-shot or languageconditioned behavior, but benchmarks are still dominated by simulation. In this work, we focus on the fine-grained behavior of fast-moving real robots and present a largescale experimental study involving 262 navigation episodes in a real environment with a physical robot, where we analyze the type of reasoning emerging from end-to-end training. In particular, we study the presence of realistic dynamics which the agent learned for open-loop forecasting, and their interplay with sensing. We analyze the way the agent uses latent memory to hold elements of the scene structure and information gathered during exploration. We probe the planning capabilities of the agent, and find in its memory evidence for somewhat precise plans over a limited horizon. Furthermore, we show in a post-hoc analysis that the value function learned by the agent relates to long-term planning. Put together, our experiments paint a new picture on how using tools from computer vision and sequential decision making have led to new capabilities in robotics and control. An interactive tool is available [here].
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 474db206-b05d-45d6-a722-58cb9218f7efCited by top-tier papers2
- Distilling LLM Prior to Flow Model for Generalizable Agent's Imagination in Object Goal NavigationBadi Li, Renjie Lu, Yu Zhou, Jingke Meng et al.NeurIPS 2025 · 5 citations
- Expand Your SCOPE: Semantic Cognition over Potential-Based Exploration for Embodied Visual NavigationNingnan Wang, Weihuang Chen, Liming Chen, Haoxuan Ji et al.AAAI 2026
Builds on20
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- Learning Mesh-Based Simulation with Graph NetworksTobias Pfaff, Meire Fortunato, Alvaro Sanchez-Gonzalez, Peter W. BattagliaICLR 2021 · 1,175 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Object Goal Navigation using Goal-Oriented Semantic ExplorationDevendra Singh Chaplot, Dhiraj Gandhi, Abhinav Gupta, Ruslan SalakhutdinovNeurIPS 2020 · 857 citations
- Learning To Explore Using Active Neural SLAMDevendra Singh Chaplot, Dhiraj Gandhi, Saurabh Gupta, Abhinav Gupta et al.ICLR 2020 · 603 citations
Related papers
- Learning to Navigate Efficiently and Precisely in Real EnvironmentsGuillaume Bono, Hervé Poirier, Leonid Antsfeld, Gianluca Monaci et al.CVPR 2024
- Planning from Imagination: Episodic Simulation and Episodic Memory for Vision-and-Language NavigationYiyuan Pan, Yunzhe Xu, Zhe Liu, Hesheng WangAAAI 2025 · 8 citations
- Robot-R1: Reinforcement Learning for Enhanced Embodied Reasoning in RoboticsDongyoung Kim, Sumin Park, Huiwon Jang, Jinwoo Shin et al.NeurIPS 2025 · 29 citations
- Grounded Decoding: Guiding Text Generation with Grounded Models for Embodied AgentsWenlong Huang, Fei Xia, Dhruv Shah, Danny Driess et al.NeurIPS 2023 · 102 citations
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsGengze Zhou, Yicong Hong, Qi WuAAAI 2024 · 361 citations
