Evolving Graphical Planner: Contextual Global Planning for Vision-and-Language Navigation
Zhiwei Deng, Karthik Narasimhan, Olga Russakovsky
Abstract
The ability to perform effective planning is crucial for building an instructionfollowing agent. When navigating through a new environment, an agent is challenged with (1) connecting the natural language instructions with its progressively growing knowledge of the world; and (2) performing long-range planning and decision making in the form of effective exploration and error correction. Current methods are still limited on both fronts despite extensive efforts. In this paper, we introduce the Evolving Graphical Planner (EGP), a model that performs global planning for navigation based on raw sensory input. The model dynamically constructs a graphical representation, generalizes the action space to allow for more flexible decision making, and performs efficient planning on a proxy graph representation. We evaluate our model on a challenging Vision-and-Language Navigation (VLN) task with photorealistic images, and achieve superior performance compared to previous navigation architectures. For instance, we achieve a 53% success rate on the test split of the Room-to-Room navigation task [1] through pure imitation learning, outperforming previous navigation architectures by up to 5%. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers45
- History Aware Multimodal Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, Ivan LaptevNeurIPS 2021 · 427 citations
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsGengze Zhou, Yicong Hong, Qi WuAAAI 2024 · 361 citations
- Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi, Cordelia Schmid et al.CVPR 2022 · 150 citations
- Scaling Data Generation in Vision-and-Language NavigationZun Wang, Jialu Li, Yicong Hong, Yi Wang et al.ICCV 2023 · 136 citations
- Vision-Language Navigation with Random Environmental MixupChong Liu, Fengda Zhu, Xiaojun Chang, Xiaodan Liang et al.ICCV 2021 · 113 citations
Builds on4
- Sparse Graphical Memory for Robust PlanningScott Emmons, Ajay Jain, Michael Laskin, Thanard Kurutach et al.NeurIPS 2020 · 60 citations
- Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-TrainingWeituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin et al.CVPR 2020
- Spatio-Temporal Graph for Video Captioning With Knowledge DistillationBoxiao Pan, Haoye Cai, De-An Huang, Kuan-Hui Lee et al.CVPR 2020
- Vision-Language Navigation With Self-Supervised Auxiliary Reasoning TasksFengda Zhu, Yi Zhu, Xiaojun Chang, Xiaodan LiangCVPR 2020
Related papers
- AwareVLN: Reasoning with Self-awareness for Vision-Language NavigationWenxuan Guo, Xiuwei Xu, Yichen Liu, Xiangyu Li et al.CVPR 2026 · 7 citations
- Target-Driven Structured Transformer Planner for Vision-Language NavigationYusheng Zhao, Jinyu Chen, Chen Gao, Wenguan Wang et al.ACM MM 2022 · 49 citations
- SeqWalker: Sequential-Horizon Vision-and-Language Navigation with Hierarchical PlanningZebin Han, Xudong Wang, Baichen Liu, Qi Lyu et al.AAAI 2026 · 2 citations
- NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous EnvironmentsXuan Yao, Junyu Gao, Changsheng XuICCV 2025 · 7 citations
- NavForesee: A Unified Vision-Language World Model for Hierarchical Planning and Dual-Horizon Navigation PredictionFei Liu, Shichao Xie, Minghua Luo, Zedong Chu et al.CVPR 2026 · 16 citations
