Scene Map-based Prompt Tuning for Navigation Instruction Generation
Sheng Fan, Rui Liu, Wenguan Wang, Yi Yang
Abstract
Navigation instruction generation (NIG), which provides interactive feedback and guidance to humans along a trajectory, is vital for developing embodied agents capable of human-machine communication and collaboration through natural language. Early data-driven methods directly map sequences of past observations to trajectory descriptions on limited datasets, lacking the necessary spatial understanding in complex 3D environments. While recent approaches leverage Large Language Models (LLMs) to improve NIG, they often overlook the global spatial context in navigation, such as the inherent space discretization in maps. Instead of straightforwardly feeding textual descriptions of the map into LLMs, we propose a scene map-based prompt tuning framework for NIG, MAPINSTRUCTOR, which incorporates map context for parameter-efficient updating of LLMs. MAPINSTRUCTOR comprises three key components: (i) scene representation encoding, where egocentric observations are projected into 3D voxels for fine-grained scene understanding; (ii) map prompt tuning, which integrates a topological map representation of the entire trajectory into an LLM-based decoder; and (iii) landmark uncertainty assessment, which mitigates hallucinations in landmark predictions, thereby enhancing the reliability and coherence of instruction generation. Extensive experiments on three navigation datasets (i.e., R2R, REVERIE, RxR) confirm the generalization and effectiveness of our algorithm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a68e093e-882e-49ec-a21b-547784da5820Cited by top-tier papers4
- 3D Gaussian Map with Open-Set Semantic Grouping for Vision-Language NavigationJianzhe Gao, Rui Liu, Wenguan WangICCV 2025 · 5 citations
- Uncertainty-Aware Gaussian Map for Vision-Language NavigationJianzhe Gao, Rui Liu, Yuxuan Xu, Tongtong Cao et al.ICLR 2026 · 3 citations
- SKDream: Controllable Multi-view and 3D Generation with Arbitrary SkeletonsYuanyou Xu, Zongxin Yang, Yi YangCVPR 2025
- DiffVsgg: Diffusion-Driven Online Video Scene Graph GenerationMu Chen, Liulei Li, Wenguan Wang, Yi YangCVPR 2025
Builds on46
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
Related papers
- VPN: Visual Prompt NavigationShuo Feng, Zihan Wang, Yuchen Li, Rui Kong et al.AAAI 2026 · 2 citations
- VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street ViewRaphael Schumann, Wanrong Zhu, Weixi Feng, Tsu-Jui Fu et al.AAAI 2024 · 122 citations
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsGengze Zhou, Yicong Hong, Qi WuAAAI 2024 · 361 citations
- Visual-Language Navigation Pretraining via Prompt-based Environmental Self-explorationXiwen Liang, Fengda Zhu, Lingling Li, Hang Xu et al.ACL 2022 · 37 citations
- OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language MappingDanyang Li, Zenghui Yang, Guangpeng Qi, Songtao Pang et al.ACM MM 2025 · 2 citations
