HOP: History-and-Order Aware Pretraining for Vision-and-Language Navigation
Yanyuan Qiao, Yuankai Qi, Yicong Hong, Zheng Yu, Peng Wang, Qi Wu
Abstract
Pretraining has been adopted in a few of recent works for Vision-and-Language Navigation (VLN). However, pre-vious pre-training methods for VLN either lack the ability to predict future actions or ignore the trajectory contexts, which are essential for a greedy navigation process. In this work, to promote the learning of spatio-temporal visual-textual correspondence as well as the agent's capability of decision making, we propose a novel history-and-order aware pre-training paradigm (HOP) with VLN-specific objectives that exploit the past observations and support future action prediction. Specifically, in addition to the commonly used Masked Language Modeling (MLM) and Trajectory-Instruction Matching (TIM), we design two proxy tasks to model temporal order information: Trajectory Order Modeling (TOM) and Group Order Modeling (GOM). Moreover, our navigation action prediction is also enhanced by intro-ducing the task of Action Prediction with History (APH), which takes into account the history visual perceptions. Extensive experimental results on four downstream VLN tasks (R2R, REVERIE, NDH, RxR) demonstrate the effectiveness of our proposed method compared against several state-of-the-art agents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8282c763-add0-470f-bfcb-6bc2699f004aCited by top-tier papers47
- Scaling Data Generation in Vision-and-Language NavigationZun Wang, Jialu Li, Yicong Hong, Yi Wang et al.ICCV 2023 · 136 citations
- GridMM: Grid Memory Map for Vision-and-Language NavigationZihan Wang, Xiangyang Li, Jiahao Yang, Yeqi Liu et al.ICCV 2023 · 136 citations
- PanoGen: Text-Conditioned Panoramic Environment Generation for Vision-and-Language NavigationJialu Li, Mohit BansalNeurIPS 2023 · 110 citations
- Bird's-Eye-View Scene Graph for Vision-Language NavigationRui Liu, Xiaohan Wang, Wenguan Wang, Yi YangICCV 2023 · 100 citations
- Dreamwalker: Mental Planning for Continuous Vision-Language NavigationHanqing Wang, Wei Liang, Luc Van Gool, Wenguan WangICCV 2023 · 98 citations
Builds on11
- Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingAlexander Ku, Peter Anderson, Roma Patel, Eugene Ie et al.EMNLP 2020 · 208 citations
- Airbert: In-domain Pretraining for Vision-and-Language NavigationPierre-Louis Guhur, Makarand Tapaswi, Shizhe Chen, Ivan Laptev et al.ICCV 2021 · 185 citations
- Language and Visual Entity Relationship Graph for Agent NavigationYicong Hong, Cristian Rodriguez Opazo, Yuankai Qi, Qi Wu et al.NeurIPS 2020 · 167 citations
- Transferable Representation Learning in Vision-and-Language NavigationHaoshuo Huang, Vihan Jain, Harsh Mehta, Alexander Ku et al.ICCV 2019 · 93 citations
- The Road to Know-Where: An Object-and-Room Informed Sequential BERT for Indoor Vision-Language NavigationYuankai Qi, Zizheng Pan, Yicong Hong, Ming-Hsuan Yang et al.ICCV 2021 · 87 citations
Related papers
- History Aware Multimodal Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, Ivan LaptevNeurIPS 2021 · 427 citations
- Improving Vision-and-Language Navigation by Generating Future-View Image SemanticsJialu Li, Mohit BansalCVPR 2023
- Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-TrainingWeituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin et al.CVPR 2020
- VLN BERT: A Recurrent Vision-and-Language BERT for NavigationYicong Hong, Qi Wu, Yuankai Qi, Cristian Rodriguez Opazo et al.CVPR 2021
- Mind the Gap: Improving Success Rate of Vision-and-Language Navigation by Revisiting Oracle Success RoutesChongyang Zhao, Yuankai Qi, Qi WuACM MM 2023 · 17 citations
