VLN-ChEnv: Vision-language Navigation in Changeable Environments
Shubo Liu, Hongsheng Zhang, Qian Qiao, Qi Wu, Peng Wang
摘要
Commanding robots to do chores using natural language instructions has been a dream of us for a long time. The navigation capability, as one of the key foundational abilities to achieve this goal, has garnered significant attention in this regard. When human users instruct intelligent agent, the instructions they given sometimes exhibit slight discrepancies from navigable ones, as user's understanding of scene may not be up-to-date due to instant change of environments. This paper investigates 3 common scenarios where instructions and navigation scenes are imperfectly aligned: change of navigability, incorrect landmark references, and incorrect direction descriptions. We then propose an ImperfectVLN task and dataset for evaluating an agent's navigation performance under instruction and environment imperfectly matched conditions. Evaluation results indicate significant performance fluctuations in existing state-of-the-art models under modification scenarios including referred landmark removal and original path blockages. We also provide a series of result analyses and further insights. We aim for this new dataset to become a valuable benchmark, enhancing practical VLN tasks. We further design a reflection module based on our insights, allowing an agent to review its history and identify potential errors. Experiments show that this module improves the performance on ImperfectVLN by 4.4%.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Navigating Beyond Instructions: Vision-and-Language Navigation in Obstructed EnvironmentsHaodong Hong, Sen Wang, Zi Huang, Qi Wu 等ACM MM 2024 · 被引用 4 次
- REVERIE: Remote Embodied Visual Referring Expression in Real Indoor EnvironmentsYuankai Qi, Qi Wu, Peter Anderson, Xin Wang 等CVPR 2020
- VLN-Trans: Translator for the Vision and Language Navigation AgentYue Zhang, Parisa KordjamshidiACL 2023 · 被引用 6 次
- Counterfactual Cycle-Consistent Learning for Instruction Following and Generation in Vision-Language NavigationHanqing Wang, Wei Liang, Jianbing Shen, Luc Van Gool 等CVPR 2022 · 被引用 49 次
- Landmark-RxR: Solving Vision-and-Language Navigation with Fine-Grained Alignment SupervisionKeji He, Yan Huang, Qi Wu, Jianhua Yang 等NeurIPS 2021 · 被引用 55 次
