All-day Multi-scenes Lifelong Vision-and-Language Navigation with Tucker Adaptation
Xudong Wang, Gan Li, Zhiyu Liu, Yao Wang, Lianqing Liu, Zhi Han
Abstract
Deploying vision-and-language navigation (VLN) agents requires adaptation across diverse scenes and environments, but fine-tuning on a specific scenario often causes catastrophic forgetting in others, which severely limits flexible long-term deployment. We formalize this challenge as the all-day multi-scenes lifelong VLN (AML-VLN) problem. Existing parameter-efficient adapters (e.g., LoRA and its variants) are limited by their two-dimensional matrix form, which fails to capture the multi-hierarchical navigation knowledge spanning multiple scenes and environments. To address this, we propose Tucker Adaptation (TuKA), which represents the multi-hierarchical navigation knowledge as a high-order tensor and leverages Tucker decomposition to decouple the knowledge into shared subspaces and scenario-specific experts. We further introduce a decoupled knowledge incremental learning strategy to consolidate shared subspaces while constraining specific experts for decoupled lifelong learning. Building on TuKA, we also develop a VLN agent named AlldayWalker, which continually learns across multiple navigation scenarios, achieving all-day multi-scenes navigation. Extensive experiments show that AlldayWalker consistently outperforms state-of-the-art baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02fb4e5d-0a40-43fd-948f-d40b9e5ee86eCited by top-tier papers3
- PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action ModelWenqi Liang, Gan Sun, Yao He, Jiahua Dong et al.ICLR 2026 · 20 citations
- Lifelong Embodied Navigation LearningXudong Wang, Jiahua Dong, Baichen Liu, Qi Lyu et al.ICLR 2026 · 5 citations
- Lifelong Language-Conditioned Robotic Manipulation LearningXudong Wang, Zebin Han, Zhiyu Liu, Gan Li et al.AAAI 2026
Builds on38
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- History Aware Multimodal Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, Ivan LaptevNeurIPS 2021 · 427 citations
- Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingAlexander Ku, Peter Anderson, Roma Patel, Eugene Ie et al.EMNLP 2020 · 208 citations
- Incremental Learning Using Conditional Adversarial NetworksYe Xiang, Ying Fu, Pan Ji, Hua HuangICCV 2019 · 188 citations
- Language Models Meet World Models: Embodied Experiences Enhance Language ModelsJiannan Xiang, Tianhua Tao, Yi Gu, Tianmin Shu et al.NeurIPS 2023 · 180 citations
Related papers
- ME: Continual Vision-and-Language Navigation via Mixture of Macro and Micro ExpertsYongliang Jiang, Huaidong Zhang, Xuandi Luo, Shengfeng HeICLR 2026
- TuckA: Hierarchical Compact Tensor Experts for Efficient Fine-TuningQifeng Lei, Zhiyong Yang, Qianqian Xu, Cong Hua et al.AAAI 2026
- Hierarchical-Task-Aware Multi-modal Mixture of Incremental LoRA Experts for Embodied Continual LearningZiqi Jia, Anmin Wang, Xiaoyang Qu, Xiaowen Yang et al.ACL 2025
- SeqWalker: Sequential-Horizon Vision-and-Language Navigation with Hierarchical PlanningZebin Han, Xudong Wang, Baichen Liu, Qi Lyu et al.AAAI 2026 · 2 citations
- Turning Adaptation into Assets: Cross-Domain Bridging for Online Vision-Language NavigationZixuan Hu, Xuantuo Huang, Yancheng Li, Yichun Hu et al.ICML 2026
