VLN-PETL: Parameter-Efficient Transfer Learning for Vision-and-Language Navigation
Yanyuan Qiao, Zheng Yu, Qi Wu
Abstract
The performance of the Vision-and-Language Navigation (VLN) tasks has witnessed rapid progress recently thanks to the use of large pre-trained vision-and-language models. However, full fine-tuning the pre-trained model for every downstream VLN task is becoming costly due to the considerable model size. Recent research hotspot of Parameter-Efficient Transfer Learning (PETL) shows great potential in efficiently tuning large pre-trained models for the common CV and NLP tasks, which exploits the most of the representation knowledge implied in the pre-trained model while only tunes a minimal set of parameters. However, simply utilizing existing PETL methods for the more challenging VLN tasks may bring non-trivial degeneration to the performance. Therefore, we present the first study to explore PETL methods for VLN tasks and propose a VLNspecific PETL method named VLN-PETL. Specifically, we design two PETL modules: Historical Interaction Booster (HIB) and Cross-modal Interaction Booster (CIB). Then we combine these two modules with several existing PETL methods as the integrated VLN-PETL. Extensive experimental results on four mainstream VLN tasks (R2R, REVERIE, NDH, RxR) demonstrate the effectiveness of our proposed VLN-PETL, where VLN-PETL achieves comparable or even better performance to full fine-tuning and outperforms other PETL methods with promising margins. The source code is available at https://github.com/YanyuanQiao/VLN-PETL * Corresponding author Updated Frozen "Go straight past the bed and the fireplace. Exit the room using the door on the left of the fireplace. Wait there.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Towards Learning a Generalist Model for Embodied NavigationDuo Zheng, Shijia Huang, Lin Zhao, Yiwu Zhong et al.CVPR 2024 · 37 citations
- Vision-and-Language Navigation via Causal LearningLiuyi Wang, Zongtao He, Ronghao Dang, Mengjiao Shen et al.CVPR 2024 · 24 citations
- Fast-Slow Test-Time Adaptation for Online Vision-and-Language NavigationJunyu Gao, Xuan Yao, Changsheng XuICML 2024 · 22 citations
- Navigating Beyond Instructions: Vision-and-Language Navigation in Obstructed EnvironmentsHaodong Hong, Sen Wang, Zi Huang, Qi Wu et al.ACM MM 2024 · 4 citations
- COSMO: Combination of Selective Memorization for Low-Cost Vision-and-Language NavigationSiqi Zhang, Yanyuan Qiao, Qunbo Wang, Zike Yan et al.ICCV 2025 · 3 citations
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick et al.ICLR 2022 · 1,182 citations
- Compacter: Efficient Low-Rank Hypercomplex Adapter LayersRabeeh Karimi Mahabadi, James Henderson, Sebastian RuderNeurIPS 2021 · 700 citations
Related papers
- BarLeRIa: An Efficient Tuning Framework for Referring Image SegmentationYaoming Wang, Jin Li, Xiaopeng Zhang, Bowen Shi et al.ICLR 2024 · 11 citations
- Parameter and Computation Efficient Transfer Learning for Vision-Language Pre-trained ModelsQiong Wu, Wei Yu, Yiyi Zhou, Shubin Huang et al.NeurIPS 2023 · 16 citations
- MaPPER: Multimodal Prior-guided Parameter Efficient Tuning for Referring Expression ComprehensionTing Liu, Zunnan Xu, Yue Hu, Liangtao Shi et al.EMNLP 2024 · 6 citations
- Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-TrainingWeituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin et al.CVPR 2020
- One Network, Many Masks: Towards More Parameter-Efficient Transfer LearningGuangtao Zeng, Peiyuan Zhang, Wei LuACL 2023 · 8 citations
