VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street View
Raphael Schumann, Wanrong Zhu, Weixi Feng, Tsu-Jui Fu, Stefan Riezler, William Yang Wang
Abstract
Incremental decision making in real-world environments is one of the most challenging tasks in embodied artificial intelligence. One particularly demanding scenario is Vision and Language Navigation (VLN) which requires visual and natural language understanding as well as spatial and temporal reasoning capabilities. The embodied agent needs to ground its understanding of navigation instructions in observations of a real-world environment like Street View. Despite the impressive results of LLMs in other research areas, it is an ongoing problem of how to best connect them with an interactive visual environment. In this work, we propose VELMA, an embodied LLM agent that uses a verbalization of the trajectory and of visual environment observations as contextual prompt for the next action. Visual information is verbalized by a pipeline that extracts landmarks from the human written navigation instructions and uses CLIP to determine their visibility in the current panorama view. We show that VELMA is able to successfully follow navigation instructions in Street View with only two in-context examples. We further finetune the LLM agent on a few thousand examples and achieve around 25% relative improvement in task completion over the previous state-of-the-art for two datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c581dd2f-ef12-4c96-bc14-3669b0497ad4Cited by top-tier papers15
- VirtuWander: Enhancing Multi-modal Interaction for Virtual Tour Guidance through Large Language ModelsZhan Wang, Linping Yuan, Liangwei Wang, Bingchuan Jiang et al.CHI 2024 · 74 citations
- UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban ScenariosBaichuan Zhou, Haote Yang, Dairong Chen, Junyan Ye et al.AAAI 2025 · 34 citations
- MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language NavigationJiaqi Chen, Bingqian Lin, Ran Xu, Zhenhua Chai et al.ACL 2024 · 29 citations
- Exploring the Robustness of Decision-Level Through Adversarial Attacks on LLM-Based Embodied ModelsShuyuan Liu, Jiawei Chen, Shouwei Ruan, Hang Su et al.ACM MM 2024 · 22 citations
- CityNavAgent: Aerial Vision-and-Language Navigation with Hierarchical Semantic Planning and Global MemoryWeichen Zhang, Chen Gao, Shiquan Yu, Ruiying Peng et al.ACL 2025 · 22 citations
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- History Aware Multimodal Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, Ivan LaptevNeurIPS 2021 · 427 citations
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsGengze Zhou, Yicong Hong, Qi WuAAAI 2024 · 361 citations
- ESC: Exploration with Soft Commonsense Constraints for Zero-shot Object NavigationKaiwen Zhou, Kaizhi Zheng, Connor Pryor, Yilin Shen et al.ICML 2023 · 221 citations
Related papers
- VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation AgentsXunyi Zhao, Gengze Zhou, Qi WuACL 2026 · 3 citations
- OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language MappingDanyang Li, Zenghui Yang, Guangpeng Qi, Songtao Pang et al.ACM MM 2025 · 2 citations
- MapNav: A Novel Memory Representation via Annotated Semantic Maps for VLM-based Vision-and-Language NavigationLingfeng Zhang, Xiaoshuai Hao, Qinwen Xu, Qiang Zhang et al.ACL 2025 · 55 citations
- KERM: Knowledge Enhanced Reasoning for Vision-and-Language NavigationXiangyang Li, Zihan Wang, Jiahao Yang, Yaowei Wang et al.CVPR 2023
- Scene Map-based Prompt Tuning for Navigation Instruction GenerationSheng Fan, Rui Liu, Wenguan Wang, Yi YangCVPR 2025
