NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models
Gengze Zhou, Yicong Hong, Qi Wu
Abstract
Trained with an unprecedented scale of data, large language models (LLMs) like ChatGPT and GPT-4 exhibit the emergence of significant reasoning abilities from model scaling. Such a trend underscored the potential of training LLMs with unlimited language data, advancing the development of a universal embodied agent. In this work, we introduce the NavGPT, a purely LLM-based instruction-following navigation agent, to reveal the reasoning capability of GPT models in complex embodied scenes by performing zero-shot sequential action prediction for vision-and-language navigation (VLN). At each step, NavGPT takes the textual descriptions of visual observations, navigation history, and future explorable directions as inputs to reason the agent's current status, and makes the decision to approach the target. Through comprehensive experiments, we demonstrate NavGPT can explicitly perform high-level planning for navigation, including decomposing instruction into sub-goals, integrating commonsense knowledge relevant to navigation task resolution, identifying landmarks from observed scenes, tracking navigation progress, and adapting to exceptions with plan adjustment. Furthermore, we show that LLMs is capable of generating high-quality navigational instructions from observations and actions along a path, as well as drawing accurate top-down metric trajectory given the agent's navigation history. Despite the performance of using NavGPT to zero-shot R2R tasks still falling short of trained models, we suggest adapting multi-modality inputs for LLMs to use as visual navigation agents and applying the explicit reasoning of LLMs to benefit learning-based models. Code is available at: https://github.com/GengzeZhou/NavGPT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 97cc530b-e6e2-46b0-b502-68c827b1bc9fCited by top-tier papers93
- SG-Nav: Online 3D Scene Graph Prompting for LLM-based Zero-shot Object NavigationHang Yin, Xiuwei Xu, Zhenyu Wu, Jie Zhou et al.NeurIPS 2024 · 215 citations
- Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry PriorsDuo Zheng, Shijia Huang, Yanyang Li, Liwei WangNeurIPS 2025 · 130 citations
- VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street ViewRaphael Schumann, Wanrong Zhu, Weixi Feng, Tsu-Jui Fu et al.AAAI 2024 · 122 citations
- Embodied Navigation Foundation ModelJiazhao Zhang, Anqi Li, Yunpeng Qi, Minghan Li et al.ICLR 2026 · 93 citations
- Affordances-Oriented Planning Using Foundation Models for Continuous Vision-Language NavigationJiaqi Chen, Bingqian Lin, Xinmin Liu, Lin Ma et al.AAAI 2025 · 61 citations
Builds on35
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
Related papers
- MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language NavigationJiaqi Chen, Bingqian Lin, Ran Xu, Zhenhua Chai et al.ACL 2024 · 29 citations
- VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation AgentsXunyi Zhao, Gengze Zhou, Qi WuACL 2026 · 3 citations
- NavBench: Probing Multimodal Large Language Models for Embodied NavigationYanyuan Qiao, Haodong Hong, Wenqi Lyu, Dong An et al.NeurIPS 2025 · 27 citations
- Towards Learning a Generalist Model for Embodied NavigationDuo Zheng, Shijia Huang, Lin Zhao, Yiwu Zhong et al.CVPR 2024 · 37 citations
- An Embodied Generalist Agent in 3D WorldJiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu et al.ICML 2024 · 361 citations
