Vision-and-Language Navigation via Causal Learning
Liuyi Wang, Zongtao He, Ronghao Dang, Mengjiao Shen, Chengju Liu, Qijun Chen
Abstract
In the pursuit of robust and generalizable environment perception and language understanding, the ubiquitous challenge of dataset bias continues to plague vision-andlanguage navigation (VLN) agents, hindering their performance in unseen environments. This paper introduces the generalized cross-modal causal transformer (GOAT), a pioneering solution rooted in the paradigm of causal inference. By delving into both observable and unobservable confounders within vision, language, and history, we propose the back-door and front-door adjustment causal learning (BACL and FACL) modules to promote unbiased learning by comprehensively mitigating potential spurious correlations. Additionally, to capture global confounder features, we propose a cross-modal feature pooling (CFP) module supervised by contrastive learning, which is also shown to be effective in improving cross-modal representations during pre-training. Extensive experiments across multiple VLN datasets (R2R, REVERIE, RxR, and SOON) underscore the superiority of our proposed method over previous state-of-the-art approaches. Code is available at https: //github.com/CrystalSixone/VLN-GOAT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9d4a8352-010e-464d-97cc-6539c0698eeaCited by top-tier papers21
- Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language NavigationZihan Wang, Seungjun Lee, Gim Hee LeeNeurIPS 2025 · 36 citations
- Active Test-time Vision-Language NavigationHeeju Ko, Sung June Kim, Gyeongrok Oh, Jeongyoon Yoon et al.NeurIPS 2025 · 10 citations
- Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual DisparitiesLiuyi Wang, Xinyuan Xia, Hui Zhao, Hanqing Wang et al.ICCV 2025 · 5 citations
- Revealing Multimodal Causality with Large Language ModelsJin Li, Shoujin Wang, Qi Zhang, Feng Liu et al.NeurIPS 2025 · 5 citations
- SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of ExpertsGengze Zhou, Yicong Hong, Zun Wang, Chongyang Zhao et al.ICCV 2025 · 4 citations
Builds on44
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- How Much Can CLIP Benefit Vision-and-Language Tasks?Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal et al.ICLR 2022 · 503 citations
- History Aware Multimodal Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, Ivan LaptevNeurIPS 2021 · 427 citations
Related papers
- Causal Attention for Vision-Language TasksXu Yang, Hanwang Zhang, Guojun Qi, Jianfei CaiCVPR 2021
- Contrastive Instruction-Trajectory Learning for Vision-Language NavigationXiwen Liang, Fengda Zhu, Yi Zhu, Bingqian Lin et al.AAAI 2022 · 29 citations
- Multimodal Causal Reasoning for UAV Object DetectionNianxin Li, Mao Ye, Lihua Zhou, Shuaifeng Li et al.NeurIPS 2025 · 1 citation
- ADAPT: Vision-Language Navigation with Modality-Aligned Action PromptsBingqian Lin, Yi Zhu, Zicong Chen, Xiwen Liang et al.CVPR 2022 · 45 citations
- Mind the Gap: Improving Success Rate of Vision-and-Language Navigation by Revisiting Oracle Success RoutesChongyang Zhao, Yuankai Qi, Qi WuACM MM 2023 · 17 citations
