Improving Vision-and-Language Navigation by Generating Future-View Image Semantics
Jialu Li, Mohit Bansal
摘要
Vision-and-Language Navigation (VLN) is the task that requires an agent to navigate through the environment based on natural language instructions. At each step, the agent takes the next action by selecting from a set of navigable locations. In this paper, we aim to take one step further and explore whether the agent can benefit from generating the potential future view during navigation. Intuitively, humans will have an expectation of how the future environment will look like, based on the natural language instructions and surrounding views, which will aid correct navigation. Hence, to equip the agent with this ability to generate the semantics of future navigation views, we first propose three proxy tasks during the agent's in-domain pre-training: Masked Panorama Modeling (MPM), Masked Trajectory Modeling (MTM), and Action Prediction with Image Generation (APIG). These three objectives teach the model to predict missing views in a panorama (MPM), predict missing steps in the full trajectory (MTM), and generate the next view based on the full instruction and navigation history (APIG), respectively. We then fine-tune the agent on the VLN task with an auxiliary loss that minimizes the difference between the view semantics generated by the agent and the ground truth view semantics of the next step. Empirically, our VLN-SIG achieves the new state-of-the-art on both Room-to-Room dataset and CVDN dataset. We further show that our agent learns to fill in missing patches in future views qualitatively, which brings more interpretability over agents' predicted actions. Lastly, we demonstrate that learning to predict future view semantics also enables the agent to have better performance on longer paths. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- Scaling Data Generation in Vision-and-Language NavigationZun Wang, Jialu Li, Yicong Hong, Yi Wang 等ICCV 2023 · 被引用 136 次
- PanoGen: Text-Conditioned Panoramic Environment Generation for Vision-and-Language NavigationJialu Li, Mohit BansalNeurIPS 2023 · 被引用 110 次
- Towards Learning a Generalist Model for Embodied NavigationDuo Zheng, Shijia Huang, Lin Zhao, Yiwu Zhong 等CVPR 2024 · 被引用 37 次
- Vision-and-Language Navigation via Causal LearningLiuyi Wang, Zongtao He, Ronghao Dang, Mengjiao Shen 等CVPR 2024 · 被引用 24 次
- Fast-Slow Test-Time Adaptation for Online Vision-and-Language NavigationJunyu Gao, Xuan Yao, Changsheng XuICML 2024 · 被引用 22 次
它引用的顶会 Paper20
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- History Aware Multimodal Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, Ivan LaptevNeurIPS 2021 · 被引用 427 次
- Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingAlexander Ku, Peter Anderson, Roma Patel, Eugene Ie 等EMNLP 2020 · 被引用 208 次
- Airbert: In-domain Pretraining for Vision-and-Language NavigationPierre-Louis Guhur, Makarand Tapaswi, Shizhe Chen, Ivan Laptev 等ICCV 2021 · 被引用 185 次
相关 Paper
- Vision-Language Navigation With Self-Supervised Auxiliary Reasoning TasksFengda Zhu, Yi Zhu, Xiaojun Chang, Xiaodan LiangCVPR 2020
- Generative Language-Grounded Policy in Vision-and-Language Navigation with Bayes' RuleShuhei Kurita, Kyunghyun ChoICLR 2021 · 被引用 29 次
- Lookahead Exploration with Neural Radiance Representation for Continuous Vision-Language NavigationZihan Wang, Xiangyang Li, Jiahao Yang, Yeqi Liu 等CVPR 2024 · 被引用 13 次
- Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-TrainingWeituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin 等CVPR 2020
- HOP: History-and-Order Aware Pretraining for Vision-and-Language NavigationYanyuan Qiao, Yuankai Qi, Yicong Hong, Zheng Yu 等CVPR 2022 · 被引用 71 次
