SMoLoRa: Exploring and Defying Dual Catastrophic Forgetting in Continual Visual Instruction Tuning
Ziqi Wang, Chang Che, Qi Wang, Yangyang Li, Zenglin Shi, Meng Wang
Abstract
Visual instruction tuning (VIT) enables multimodal large language models (MLLMs) to effectively handle a wide range of vision tasks by framing them as language-based instructions. Building on this, continual visual instruction tuning (CVIT) extends the capability of MLLMs to incrementally learn new tasks, accommodating evolving functionalities. While prior work has advanced CVIT through the development of new benchmarks and approaches to mitigate catastrophic forgetting, these efforts largely follow traditional continual learning paradigms, neglecting the unique challenges specific to CVIT. We identify a dual form of catastrophic forgetting in CVIT, where MLLMs not only forget previously learned visual understanding but also experience a decline in instruction following abilities as they acquire new tasks. To address this, we introduce the Separable Mixture of Low-Rank Adaptation (SMoLoRA) framework, which employs separable routing through two distinct modules—one for visual understanding and another for instruction following. This dual-routing design enables specialized adaptation in both domains, preventing forgetting while improving performance. Furthermore, we propose a new CVIT benchmark that goes beyond existing benchmarks by additionally evaluating a model's ability to generalize to unseen tasks and handle diverse instructions across various tasks. Extensive experiments demonstrate that SMoLoRA outperforms existing methods in mitigating dual forgetting, improving generalization to unseen tasks, and ensuring robustness in following diverse instructions. Code is available at https://github.com/Minato-Zackie/SMoLoRA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c5e8f45d-824a-482c-b25a-a83937c7c6a3Cited by top-tier papers5
- FlyLoRA: Boosting Task Decoupling and Parameter Efficiency via Implicit Rank-Wise Mixture-of-ExpertsHeming Zou, Yunliang Zang, Wutong Xu, Yao Zhu et al.NeurIPS 2025 · 38 citations
- Affordance-First Decomposition for Continual Learning in Video–Language UnderstandingMengzhu xu, Hanzhi Liu, Ningkang Peng, qianyu Chen et al.CVPR 2026 · 7 citations
- Sparse Spectral LoRA: Routed Experts for Medical VLMsOmid Nejati Manzari, Hojat Asgariandehkordi, Taha Koleilat, Yiming Xiao et al.CVPR 2026 · 3 citations
- LoRA in LoRA: Towards Parameter-Efficient Architecture Expansion for Continual Visual Instruction TuningChang Che, Ziqi Wang, Pengwan Yang, Cheems Wang et al.AAAI 2026
- KSS-MoE: Knowledge Space Synergy Framework in Mixture of Experts for Continual Visual Instruction TuningLingyun Song, Ziyao Chen, Kang Pan, Xiaolin Han et al.AAAI 2026
Builds on13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- Advancing SMoE for Continuous Domain Adaptation of MLLMs: Adaptive Router and Domain-Specific LossLiang Zhang, Ziyao Lu, Fandong Meng, Hui Li et al.ACL 2025 · 3 citations
- Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMsZiqi Wang, Chang Che, Qi Wang, Hui Ma et al.CVPR 2026 · 4 citations
- HMVLM: Human Motion-Vision-Language Model via MoE LoRALei Hu, Yongjing Ye, Shihong XiaNeurIPS 2025 · 1 citation
- DeLo: Dual Decomposed Low-Rank Experts Collaboration for Continual Missing Modality LearningXiwei Liu, Yulong Li, Feilong Tang, Imran RazzakAAAI 2026
- LoRASculpt: Sculpting LoRA for Harmonizing General and Specialized Knowledge in Multimodal Large Language ModelsJian Liang, Wenke Huang, Guancheng Wan, Qu Yang et al.CVPR 2025
