PASs-MoE: Mitigating Misaligned Co-drift among Router and Experts via Pathway Activation Subspaces for Continual Learning
ZhiYan Hou, Haiyun Guo, Haokai Ma, Yandu Sun, Yonghui Yang, Jinqiao Wang
摘要
Continual instruction tuning (CIT) requires multimodal large language models (MLLMs) to adapt to a stream of tasks without forgetting prior capabilities. A common strategy is to isolate updates by routing inputs to different LoRA experts. However, existing LoRAbased Mixture-of-Experts (MoE) methods often jointly update the router and experts in an indiscriminate way, causing the router's preferences to co-drift with experts' adaptation pathways and gradually deviate from early-stage input-expert specialization. We term this as Misaligned Co-drift, which blurs expert responsibilities and exacerbates forgetting. To address this, we introduce the pathway activation subspace (PASs), a LoRA-induced subspace that reflects which low-rank pathway directions an input activates in each expert, providing a capability-aligned coordinate system for routing and preservation. Based on PASs, we propose a fixed-capacity PASs-based MoE-LoRA method with two components: PAS-guided Reweighting, which calibrates routing using each expert's pathway activation signals, and PAS-aware Rank Stabilization, which selectively stabilizes rank directions important to previous tasks. Experiments on a CIT benchmark show that our approach consistently outperforms a range of conventional continual learning baselines and MoE-LoRA variants in both accuracy and anti-forgetting without adding parameters. Our code will be released upon acceptance. MLLM 𝐵 ! 𝐴 ! 𝐵 " 𝐴 " Router LLM inference on task 𝒏
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Merge to Remember: Sharpness-Aware Isotropic Merging for Continual LearningQun Yang, Enneng Yang, Wei Chen, Li Shen 等ICML 2026
- Calibrated Knowledge Aggregation in Bayesian Mixture-of-Experts for Continual VQAMahsa Mozaffari, Hitesh Sapkota, Yu Kong, Xumin Liu 等ICML 2026
它引用的顶会 Paper17
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang 等CVPR 2022 · 被引用 635 次
- Gradient Projection Memory for Continual LearningGobinda Saha, Isha Garg, Kaushik RoyICLR 2021 · 被引用 409 次
- Continual Learning in Low-rank Orthogonal SubspacesArslan Chaudhry, Naeemullah Khan, Puneet K. Dokania, Philip H. S. TorrNeurIPS 2020 · 被引用 171 次
- Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts AdaptersJiazuo Yu, Yunzhi Zhuge, Lu Zhang, Ping Hu 等CVPR 2024 · 被引用 80 次
相关 Paper
- SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction TuningZhen-Hao Xie Xie, Jun-Tao Tang, Yu-Cheng Shi, Han-Jia Ye 等ICML 2026
- Dynamic Mixture of Curriculum LoRA Experts for Continual Multimodal Instruction TuningChendi Ge, Xin Wang, Zeyang Zhang, Hong Chen 等ICML 2025
- On Token's Dilemma: Dynamic MoE with Drift-Aware Token Assignment for Continual Learning of Large Vision Language ModelsChongyang Zhao, Mingsong Li, Haodong Lu, Dong GongCVPR 2026 · 被引用 3 次
- From Experts to Bases: Orthogonal Subspace Mixture for Continual Multimodal Instruction TuningPei Chen, Xilai Wang, Qixu Shi, Zejian Li 等ACL 2026
- Advancing SMoE for Continuous Domain Adaptation of MLLMs: Adaptive Router and Domain-Specific LossLiang Zhang, Ziyao Lu, Fandong Meng, Hui Li 等ACL 2025 · 被引用 3 次
