Counterfactual Vision-and-Language Navigation: Unravelling the Unseen
Amin Parvaneh, Ehsan Abbasnejad, Damien Teney, Qinfeng Shi, Anton van den Hengel
Abstract
The task of vision-and-language navigation (VLN) requires an agent to follow text instructions to find its way through simulated household environments. A prominent challenge is to train an agent capable of generalising to new environments at test time, rather than one that simply memorises trajectories and visual details observed during training. We propose a new learning strategy that learns both from observations and generated counterfactual environments. We describe an effective algorithm to generate counterfactual observations on the fly for VLN, as linear combinations of existing environments. Simultaneously, we encourage the agent's actions to remain stable between original and counterfactual environments through our novel training objective -effectively removing spurious features that would otherwise bias the agent. Our experiments show that this technique provides significant improvements in generalisation on benchmarks for Room-to-Room navigation and Embodied Question Answering.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 913da6e5-fb33-45b5-8e78-e70139eec5bbCited by top-tier papers8
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsGengze Zhou, Yicong Hong, Qi WuAAAI 2024 · 361 citations
- Active Learning by Feature MixingAmin Parvaneh, Ehsan Abbasnejad, Damien Teney, Reza Haffari et al.CVPR 2022 · 113 citations
- Learning to Imagine: Integrating Counterfactual Thinking in Neural Discrete ReasoningMoxin Li, Fuli Feng, Hanwang Zhang, Xiangnan He et al.ACL 2022 · 39 citations
- Frequency-Enhanced Data Augmentation for Vision-and-Language NavigationKeji He, Chenyang Si, Zhihe Lu, Yan Huang et al.NeurIPS 2023 · 32 citations
- Navigating Beyond Instructions: Vision-and-Language Navigation in Obstructed EnvironmentsHaodong Hong, Sen Wang, Zi Huang, Qi Wu et al.ACM MM 2024 · 4 citations
Builds on4
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
- Adversarial Examples Improve Image RecognitionCihang Xie, Mingxing Tan, Boqing Gong, Jiang Wang et al.CVPR 2020
- Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-TrainingWeituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin et al.CVPR 2020
- Counterfactual Vision and Language LearningEhsan Abbasnejad, Damien Teney, Amin Parvaneh, Javen Shi et al.CVPR 2020
Related papers
- Curriculum Learning for Vision-and-Language NavigationJiwen Zhang, Zhongyu Wei, Jianqing Fan, Jiajie PengNeurIPS 2021 · 33 citations
- Generative Language-Grounded Policy in Vision-and-Language Navigation with Bayes' RuleShuhei Kurita, Kyunghyun ChoICLR 2021 · 29 citations
- Counterfactual Cycle-Consistent Learning for Instruction Following and Generation in Vision-Language NavigationHanqing Wang, Wei Liang, Jianbing Shen, Luc Van Gool et al.CVPR 2022 · 49 citations
- Learning Vision-and-Language Navigation from YouTube VideosKunyang Lin, Peihao Chen, Diwei Huang, Thomas H. Li et al.ICCV 2023 · 57 citations
- Contrastive Instruction-Trajectory Learning for Vision-Language NavigationXiwen Liang, Fengda Zhu, Yi Zhu, Bingqian Lin et al.AAAI 2022 · 29 citations
