History to Future: Evolving Agent with Experience and Thought for Zero-shot Vision-and-Language Navigation
Guangzhao Dai, Shuo Wang, Zihan Wang, Guo-Sen Xie, Yang Yang, Jinshan Pan, Qianru Sun, Xiangbo Shu
摘要
Vision-and-Language Navigation in Continuous Environment (VLN-CE) requires an agent to follow language instructions to navigate the target destination. With the advancement of large language models (LLMs), recent efforts have explored adapting them for zero-shot VLN-CE, offering a promising solution in addressing the drawbacks of poor generalization in the training-based paradigm. However, existing LLM-based works primarily perform naive reasoning for decision-making and lack feedback, e.g., reviewing historical errors and predicting future potentials. Consequently, it may suffer from continuous failure for those initial error tasks. In this paper, we rethink LLM-based zero-shot VLN-CE and propose a new paradigm, named EVONAV, to improve the agent's decision-making with future thought and history experience via Future Chain-of-Thought (F-CoT) and History Chainof-Experience (H-CoE). F-CoT predicts future actions and landmarks as thoughts to assist navigation progress estimation and direction selection, while H-CoE summarizes historical trajectories and scenes as experience to improve navigation decision reliability. Both F-CoT and H-CoE cooperatively evolve the agent's decision-making. Extensive experiments in both the simulator and real-world environments demonstrate the effectiveness of our EVONAV.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsGengze Zhou, Yicong Hong, Qi WuAAAI 2024 · 被引用 361 次
- Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingAlexander Ku, Peter Anderson, Roma Patel, Eugene Ie 等EMNLP 2020 · 被引用 208 次
相关 Paper
- VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation AgentsXunyi Zhao, Gengze Zhou, Qi WuACL 2026 · 被引用 3 次
- NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous EnvironmentsXuan Yao, Junyu Gao, Changsheng XuICCV 2025 · 被引用 7 次
- ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language NavigationWei Xue, Mingcheng Li, Xuecheng Wu, Jingqun Tang 等CVPR 2026 · 被引用 4 次
- Aux-Think: Exploring Reasoning Strategies for Data-Efficient Vision-Language NavigationShuo Wang, Yongcai Wang, Wanting Li, Xudong Cai 等NeurIPS 2025 · 被引用 28 次
- Run, Ruminate, and Regulate: A Dual-process Thinking System for Vision-and-Language NavigationYu Zhong, Zihao Zhang, Rui Zhang, Lingdong Huang 等AAAI 2026
