Ali-UI: Enhancing Complex Vision-Language Navigation with Alignment of Unified Map and Instruction Parsing
Shanshan Li, Jiawei Hou, Da Huang, Yanwei Fu, Xiangyang Xue
Abstract
Visual language navigation (VLN) poses challenges in guiding agents through unseen environments based on natural language instructions. Existing methods either rely on imitation learning, for which training across various complex scenarios remains challenging, or leverage large visual language models (LVLMs) for zero-shot object recognition and expert iterative reasoning for improved scene understanding. Although LVLMs enhance target detection generalization, current VLN methods lack robustness in terms of environmental generalization and struggle with multi-step, coarsely directed instructions. Addressing these challenges, we introduce Ali-UI, a novel vision-language navigation approach that enables agents to navigate from random starting points in unvisited scenes and handle complex multi-step instructions. Specifically, we incorporate continuously accumulating global grid maps and local semantic maps as scene memory by employing frontier-based exploration. Multi-step coarsely directed commands are broken down with the assistance of LLaVA and matched with the scene, considering temporal and spatial alignment. Panoramic data are saved in topological form and queried by instruction segments for sequential navigation. Extensive experiments carried out in simulated environments demonstrate that Ali-UI outperforms existing state-of-the-art methods in terms of flexible human instructions and scene generalization, with the success rate improved by 23.37% and the SPL increased by 19.27% in R2R dataset.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 3fb187e3-89b3-4489-9942-e73a627ad4ecRelated papers
- RILA: Reflective and Imaginative Language Agent for Zero-Shot Semantic Audio-Visual NavigationZeyuan Yang, Jiageng Lin, Peihao Chen, Anoop Cherian et al.CVPR 2024 · 5 citations
- SeqWalker: Sequential-Horizon Vision-and-Language Navigation with Hierarchical PlanningZebin Han, Xudong Wang, Baichen Liu, Qi Lyu et al.AAAI 2026 · 2 citations
- GridMM: Grid Memory Map for Vision-and-Language NavigationZihan Wang, Xiangyang Li, Jiahao Yang, Yeqi Liu et al.ICCV 2023 · 136 citations
- Structured Scene Memory for Vision-Language NavigationHanqing Wang, Wenguan Wang, Wei Liang, Caiming Xiong et al.CVPR 2021
- OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language MappingDanyang Li, Zenghui Yang, Guangpeng Qi, Songtao Pang et al.ACM MM 2025 · 2 citations
