History Aware Multimodal Transformer for Vision-and-Language Navigation
Shizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, Ivan Laptev
Abstract
Vision-and-language navigation (VLN) aims to build autonomous visual agents that follow instructions and navigate in real scenes. To remember previously visited locations and actions taken, most approaches to VLN implement memory using recurrent states. Instead, we introduce a History Aware Multimodal Transformer (HAMT) to incorporate a long-horizon history into multimodal decision making. HAMT efficiently encodes all the past panoramic observations via a hierarchical vision transformer (ViT), which first encodes individual images with ViT, then models spatial relation between images in a panoramic observation and finally takes into account temporal relation between panoramas in the history. It, then, jointly combines text, history and current observation to predict the next action. We first train HAMT end-to-end using several proxy tasks including single step action prediction and spatial relation prediction, and then use reinforcement learning to further improve the navigation policy. HAMT achieves new state of the art on a broad range of VLN tasks, including VLN with fine-grained instructions (R2R, RxR), high-level instructions (R2R-Last, REVERIE), dialogs (CVDN) as well as long-horizon VLN (R4R, R2R-Back). We demonstrate HAMT to be particularly effective for navigation tasks with longer trajectories.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e06388eb-2442-41f3-b92a-afde563a4f93Cited by top-tier papers99
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsGengze Zhou, Yicong Hong, Qi WuAAAI 2024 · 361 citations
- Language Conditioned Spatial Relation Reasoning for 3D Object GroundingShizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi, Cordelia Schmid et al.NeurIPS 2022 · 173 citations
- Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi, Cordelia Schmid et al.CVPR 2022 · 150 citations
- Scaling Data Generation in Vision-and-Language NavigationZun Wang, Jialu Li, Yicong Hong, Yi Wang et al.ICCV 2023 · 136 citations
- GridMM: Grid Memory Map for Vision-and-Language NavigationZihan Wang, Xiangyang Li, Jiahao Yang, Yeqi Liu et al.ICCV 2023 · 136 citations
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
Related papers
- VLN BERT: A Recurrent Vision-and-Language BERT for NavigationYicong Hong, Qi Wu, Yuankai Qi, Cristian Rodriguez Opazo et al.CVPR 2021
- HOP: History-and-Order Aware Pretraining for Vision-and-Language NavigationYanyuan Qiao, Yuankai Qi, Yicong Hong, Zheng Yu et al.CVPR 2022 · 71 citations
- History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language NavigationXichen Ding, Jianzhe Gao, Cong Pan, Wenguan Wang et al.AAAI 2026 · 3 citations
- Target-Driven Structured Transformer Planner for Vision-Language NavigationYusheng Zhao, Jinyu Chen, Chen Gao, Wenguan Wang et al.ACM MM 2022 · 49 citations
- Tree-Structured Trajectory Encoding for Vision-and-Language NavigationXinzhe Zhou, Yadong MuAAAI 2023 · 2 citations
