Bridging the Gap Between Learning in Discrete and Continuous Environments for Vision-and-Language Navigation
Yicong Hong, Zun Wang, Qi Wu, Stephen Gould
摘要
Most existing works in vision-and-language navigation (VLN) focus on either discrete or continuous environments, training agents that cannot generalize across the two. Although learning to navigate in continuous spaces is closer to the real-world, training such an agent is significantly more difficult than training an agent in discrete spaces. However, recent advances in discrete VLN are challenging to translate to continuous VLN due to the domain gap. The fundamental difference between the two setups is that discrete navigation assumes prior knowledge of the connectivity graph of the environment, so that the agent can effectively transfer the problem of navigation with low-level controls to jumping from node to node with high-level actions by grounding to an image of a navigable direction. To bridge the discrete-to-continuous gap, we propose a predictor to generate a set of candidate waypoints during navigation, so that agents designed with high-level actions can be transferred to and trained in continuous environments. We refine the connectivity graph of Matterport3D to fit the continuous Habitat-Matterport3D, and train the waypoints predictor with the refined graphs to produce accessible waypoints at each time step. Moreover, we demonstrate that the predicted waypoints can be augmented during training to diversify the views and paths, and therefore enhance agent's generalization ability. Through extensive experiments we show that agents navigating in continuous environments with predicted waypoints perform significantly better than agents using low-level actions, which reduces the absolute discrete-to-continuous gap by 11.76% Success Weighted by Path Length (SPL) for the Cross-Modal Matching Agent and 18.24% SPL for the VLN$$BERT. Our agents, trained with a simple imitation learning objective, outperform previous methods by a large margin, achieving new state-of-the-art results on the testing environments of the R2R-CE and the RxR-CE datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper53
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsGengze Zhou, Yicong Hong, Qi WuAAAI 2024 · 被引用 361 次
- Weakly-Supervised Multi-Granularity Map Learning for Vision-and-Language NavigationPeihao Chen, Dongyu Ji, Kunyang Lin, Runhao Zeng 等NeurIPS 2022 · 被引用 143 次
- Scaling Data Generation in Vision-and-Language NavigationZun Wang, Jialu Li, Yicong Hong, Yi Wang 等ICCV 2023 · 被引用 136 次
- GridMM: Grid Memory Map for Vision-and-Language NavigationZihan Wang, Xiangyang Li, Jiahao Yang, Yeqi Liu 等ICCV 2023 · 被引用 136 次
- JanusVLN: Decoupling Semantics and Spatiality with Dual Implicit Memory for Vision-Language NavigationShuang Zeng, Dekang Qi, Xinyuan Chang, Feng Xiong 等ICLR 2026 · 被引用 124 次
它引用的顶会 Paper15
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion FramesErik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee 等ICLR 2020 · 被引用 608 次
- Learning To Explore Using Active Neural SLAMDevendra Singh Chaplot, Dhiraj Gandhi, Saurabh Gupta, Abhinav Gupta 等ICLR 2020 · 被引用 603 次
- Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingAlexander Ku, Peter Anderson, Roma Patel, Eugene Ie 等EMNLP 2020 · 被引用 208 次
- Language and Visual Entity Relationship Graph for Agent NavigationYicong Hong, Cristian Rodriguez Opazo, Yuankai Qi, Qi Wu 等NeurIPS 2020 · 被引用 167 次
相关 Paper
- Narrowing the Gap between Vision and Action in NavigationYue Zhang, Parisa KordjamshidiACM MM 2024 · 被引用 2 次
- Affordances-Oriented Planning Using Foundation Models for Continuous Vision-Language NavigationJiaqi Chen, Bingqian Lin, Xinmin Liu, Lin Ma 等AAAI 2025 · 被引用 61 次
- PanoGen: Text-Conditioned Panoramic Environment Generation for Vision-and-Language NavigationJialu Li, Mohit BansalNeurIPS 2023 · 被引用 110 次
- AwareVLN: Reasoning with Self-awareness for Vision-Language NavigationWenxuan Guo, Xiuwei Xu, Yichen Liu, Xiangyu Li 等CVPR 2026 · 被引用 7 次
- Learning Navigational Visual Representations with Semantic Map SupervisionYicong Hong, Yang Zhou, Ruiyi Zhang, Franck Dernoncourt 等ICCV 2023 · 被引用 56 次
