Turning Adaptation into Assets: Cross-Domain Bridging for Online Vision-Language Navigation
Zixuan Hu, Xuantuo Huang, Yancheng Li, Yichun Hu, Shengyong Xu, LINGYU DUAN
Abstract
Navigating under non-stationary environment shifts poses a critical challenge for a Vision-and-Language Navigation (VLN) agent deployed in the wild. Yet, existing Test-Time Adaptation (TTA) methods for VLN largely treat online adaptation as transient, isolated updates, leading to catastrophic forgetting and negative transfer. To overcome these issues, we propose I nter- D omain Bridg E with Historical A ssets ( IDEA ), a novel TTA framework that transforms adaptation into the accumulation and composition of assets. Specifically, IDEA introduces soft prompts optimized via a Fisher-guided weighting scheme to capture the transferable knowledge. These optimized prompts are then augmented with domain coordinates to form a dynamic asset library. Leveraging this library, IDEA constructs a cross-domain bridge by projecting the target domain onto the convex hull of historical knowledge. These designs form a complementary loop: the evolving library underpins bridge construction, while the bridge provides superior initialization to accelerate asset optimization. Extensive experiments across REVERIE, R2R, and R2R-CE benchmarks demonstrate the consistent superiority of IDEA over existing methods, showcasing its ability to enable training-free adaptation via asset sharing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6eda7ed2-e617-4e0b-9e39-b5bac5354202Builds on28
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- Test-Time Classifier Adjustment Module for Model-Agnostic Domain GeneralizationYusuke Iwasawa, Yutaka MatsuoNeurIPS 2021 · 456 citations
- History Aware Multimodal Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, Ivan LaptevNeurIPS 2021 · 427 citations
- Airbert: In-domain Pretraining for Vision-and-Language NavigationPierre-Louis Guhur, Makarand Tapaswi, Shizhe Chen, Ivan Laptev et al.ICCV 2021 · 185 citations
Related papers
- BoostAdapter: Improving Vision-Language Test-Time Adaptation via Regional BootstrappingTaolin Zhang, Jinpeng Wang, Hang Guo, Tao Dai et al.NeurIPS 2024 · 30 citations
- Decorate the Newcomers: Visual Domain Prompt for Continual Test Time AdaptationYulu Gan, Yan Bai, Yihang Lou, Xianzheng Ma et al.AAAI 2023 · 145 citations
- Test-Time Adaptation for Online Vision-Language Navigation with Feedback-based Reinforcement LearningSungjune Kim, Gyeongrok Oh, Heeju Ko, Daehyun Ji et al.ICML 2025
- ViDA: Homeostatic Visual Domain Adapter for Continual Test Time AdaptationJiaming Liu, Senqiao Yang, Peidong Jia, Renrui Zhang et al.ICLR 2024 · 71 citations
- Dance Across Shifts: Forward-Facilitation Continual Test-Time Adaptation through Dynamic Style BridgingZhilin Zhu, Yabin Wang, Zhiheng Ma, Yaguang Song et al.CVPR 2026 · 2 citations
