OVER-NAV: Elevating Iterative Vision-and-Language Navigation with Open-Vocabulary Detection and StructurEd Representation
Ganlong Zhao, Guanbin Li, Weikai Chen, Yizhou Yu
摘要
Recent advances in Iterative Vision-and-Language Navigation (IVLN) introduce a more meaningful and practical paradigm of VLN by maintaining the agent's memory across tours of scenes. Although the long-term memory aligns better with the persistent nature of the VLN task, it poses more challenges on how to utilize the highly unstructured navigation memory with extremely sparse supervision. Towards this end, we propose OVER-NAV, which aims to go over and beyond the current arts of IVLN techniques. In particular, we propose to incorporate LLMs and open-vocabulary detectors to distill key information and establish correspondence between multi-modal signals. Such a mechanism introduces reliable cross-modal supervision and enables on-the-fly generalization to unseen scenes without the need of extra annotation and re-training. To fully exploit the interpreted navigation data, we further introduce a structured representation, coded Omnigraph, to effectively integrate multi-modal information along the tour. Accompanied with a novel omnigraph fusion mechanism, OVER-NAV is able to extract the most relevant knowledge from omnigraph for a more accurate navigating action. In addition, OVER-NAV seamlessly supports both discrete and continuous environments under a unified framework. We demonstrate the superiority of OVER-NAV in extensive experiments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- REGNav: Room Expert Guided Image-Goal NavigationPengna Li, Kangyi Wu, Jingwen Fu, Sanping ZhouAAAI 2025 · 被引用 15 次
- FLARE: A Failure-Aware Framework for Autonomous Correction and Recovery in Visual-Language Robotic ManipulationGanlong Zhao, Zijia Tang, Xingping Chen, Zhanghui Kuang 等CVPR 2026 · 被引用 10 次
- PhysVLM: Enabling Visual Language Models to Understand Robotic Physical ReachabilityWeijie Zhou, Manli Tao, Chaoyang Zhao, Haiyun Guo 等CVPR 2025
- General Scene Adaptation for Vision-and-Language NavigationHaodong Hong, Yanyuan Qiao, Sen Wang, Jiajun Liu 等ICLR 2025
- Towards Open Environments and Instructions: General Vision-Language Navigation via Fast-Slow Interactive ReasoningYang Li, Aming Wu, Zihao Zhang, Yahong HanCVPR 2026
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
相关 Paper
- Iterative Vision-and-Language NavigationJacob Krantz, Shurjo Banerjee, Wang Zhu, Jason J. Corso 等CVPR 2023
- Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-TrainingWeituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin 等CVPR 2020
- Ali-UI: Enhancing Complex Vision-Language Navigation with Alignment of Unified Map and Instruction ParsingShanshan Li, Jiawei Hou, Da Huang, Yanwei Fu 等ACM MM 2025
- Cross-modal Map Learning for Vision and Language NavigationGeorgios Georgakis, Karl Schmeckpeper, Karan Wanchoo, Soham Dan 等CVPR 2022 · 被引用 2 次
- UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World ModelChangxin Huang, Lv Tang, Zhaohuan Zhan, Lisha Yu 等AAAI 2026 · 被引用 2 次
