Learning to Follow Directions in Street View
Karl Moritz Hermann, Mateusz Malinowski, Piotr Mirowski, Andras Banki-Horvath, Keith Anderson, Raia Hadsell
Abstract
Navigating and understanding the real world remains a key challenge in machine learning and inspires a great variety of research in areas such as language grounding, planning, navigation and computer vision. We propose an instruction-following task that requires all of the above, and which combines the practicality of simulated environments with the challenges of ambiguous, noisy real world data. StreetNav is built on top of Google Street View and provides visually accurate environments representing real places. Agents are given driving instructions which they must learn to interpret in order to successfully navigate in this environment. Since humans equipped with driving instructions can readily navigate in previously unseen cities, we set a high bar and test our trained agents for similar cognitive capabilities. Although deep reinforcement learning (RL) methods are frequently evaluated only on data that closely follow the training distribution, our dataset extends to multiple cities and has a clean train/test separation. This allows for thorough testing of generalisation ability. This paper presents the StreetNav environment and tasks, models that establish strong baselines, and extensive analysis of the task and the trained agents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- Waypoint Models for Instruction-guided Navigation in Continuous EnvironmentsJacob Krantz, Aaron Gokaslan, Dhruv Batra, Stefan Lee et al.ICCV 2021 · 153 citations
- Transferable Representation Learning in Vision-and-Language NavigationHaoshuo Huang, Vihan Jain, Harsh Mehta, Alexander Ku et al.ICCV 2019 · 93 citations
- Envedit: Environment Editing for Vision-and-Language NavigationJialu Li, Hao Tan, Mohit BansalCVPR 2022 · 76 citations
- Analyzing Generalization of Vision and Language Navigation to Unseen Outdoor AreasRaphael Schumann, Stefan RiezlerACL 2022 · 38 citations
- Safe Reinforcement Learning with Natural Language ConstraintsTsung-Yen Yang, Michael Y. Hu, Yinlam Chow, Peter J. Ramadge et al.NeurIPS 2021 · 37 citations
Builds on1
Related papers
- VLN-Video: Utilizing Driving Videos for Outdoor Vision-and-Language NavigationJialu Li, Aishwarya Padmakumar, Gaurav S. Sukhatme, Mohit BansalAAAI 2024 · 13 citations
- VLN-Trans: Translator for the Vision and Language Navigation AgentYue Zhang, Parisa KordjamshidiACL 2023 · 6 citations
- NaVLA: A Vision-Language-Audio-Action Model for Multimodal Instruction NavigationJugang Fan, Peihao Chen, Changhao Li, Qing Du et al.AAAI 2026
- Generating Landmark Navigation Instructions from Maps as a Graph-to-Text ProblemRaphael Schumann, Stefan RiezlerACL 2021
- UrbanNav: Learning Language-Guided Embodied Urban Navigation from Web-Scale Human TrajectoriesYanghong Mei, Yirong Yang, Longteng Guo, Qunbo Wang et al.AAAI 2026
