PASs-MoE: Mitigating Misaligned Co-drift among Router and Experts via Pathway Activation Subspaces for Continual Learning
ZhiYan Hou, Haiyun Guo, Haokai Ma, Yandu Sun, Yonghui Yang, Jinqiao Wang
Abstract
Continual instruction tuning (CIT) requires multimodal large language models (MLLMs) to adapt to a stream of tasks without forgetting prior capabilities. A common strategy is to isolate updates by routing inputs to different LoRA experts. However, existing LoRAbased Mixture-of-Experts (MoE) methods often jointly update the router and experts in an indiscriminate way, causing the router's preferences to co-drift with experts' adaptation pathways and gradually deviate from early-stage input-expert specialization. We term this as Misaligned Co-drift, which blurs expert responsibilities and exacerbates forgetting. To address this, we introduce the pathway activation subspace (PASs), a LoRA-induced subspace that reflects which low-rank pathway directions an input activates in each expert, providing a capability-aligned coordinate system for routing and preservation. Based on PASs, we propose a fixed-capacity PASs-based MoE-LoRA method with two components: PAS-guided Reweighting, which calibrates routing using each expert's pathway activation signals, and PAS-aware Rank Stabilization, which selectively stabilizes rank directions important to previous tasks. Experiments on a CIT benchmark show that our approach consistently outperforms a range of conventional continual learning baselines and MoE-LoRA variants in both accuracy and anti-forgetting without adding parameters. Our code will be released upon acceptance. MLLM 𝐵 ! 𝐴 ! 𝐵 " 𝐴 " Router LLM inference on task 𝒏
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a5b1e9ae-e843-417b-ad0b-6cca0b9b0803Cited by top-tier papers2
- Merge to Remember: Sharpness-Aware Isotropic Merging for Continual LearningQun Yang, Enneng Yang, Wei Chen, Li Shen et al.ICML 2026
- Calibrated Knowledge Aggregation in Bayesian Mixture-of-Experts for Continual VQAMahsa Mozaffari, Hitesh Sapkota, Yu Kong, Xumin Liu et al.ICML 2026
Builds on17
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang et al.CVPR 2022 · 635 citations
- Gradient Projection Memory for Continual LearningGobinda Saha, Isha Garg, Kaushik RoyICLR 2021 · 409 citations
- Continual Learning in Low-rank Orthogonal SubspacesArslan Chaudhry, Naeemullah Khan, Puneet K. Dokania, Philip H. S. TorrNeurIPS 2020 · 171 citations
- Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts AdaptersJiazuo Yu, Yunzhi Zhuge, Lu Zhang, Ping Hu et al.CVPR 2024 · 80 citations
Related papers
- SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction TuningZhen-Hao Xie Xie, Jun-Tao Tang, Yu-Cheng Shi, Han-Jia Ye et al.ICML 2026
- Dynamic Mixture of Curriculum LoRA Experts for Continual Multimodal Instruction TuningChendi Ge, Xin Wang, Zeyang Zhang, Hong Chen et al.ICML 2025
- On Token's Dilemma: Dynamic MoE with Drift-Aware Token Assignment for Continual Learning of Large Vision Language ModelsChongyang Zhao, Mingsong Li, Haodong Lu, Dong GongCVPR 2026 · 3 citations
- From Experts to Bases: Orthogonal Subspace Mixture for Continual Multimodal Instruction TuningPei Chen, Xilai Wang, Qixu Shi, Zejian Li et al.ACL 2026
- Advancing SMoE for Continuous Domain Adaptation of MLLMs: Adaptive Router and Domain-Specific LossLiang Zhang, Ziyao Lu, Fandong Meng, Hui Li et al.ACL 2025 · 3 citations
