Spectral Mixture-of-Experts for Continual Learning
Chen Yin, Xingbo Dong, Xuelin Shen, Zhe Jin
Abstract
While Parameter-Efficient Fine-Tuning using Mixtureof-Experts (MoE) is a promising solution for continual learning (CL), it suffers from two critical failure modes: structural interference, where expert updates interfere, and compositional forgetting, where the model's routing policy drifts. To address these issues, we introduce Spectral MoE, a novel framework built for CL from three core components. First, Spectral Experts are parameterized using unique, disjoint spectral masks to confine their learnable parameters to distinct frequency subspaces, ensuring a priori orthogonal updates that prevent structural interference. Second, a Dual-Router mechanism decouples online routing that learns new tasks from an offline memory that archives historical expert importance. Finally, this offline memory enables a Dynamic Consistency Projection, a geometric constraint that suppresses router drift and adaptively shields experts based on their past contributions, mitigating compositional forgetting. Validated on a strict crossdomain CL benchmark, our framework significantly outperforms existing methods, demonstrating superior knowledge retention and plasticity for new tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang et al.CVPR 2022 · 635 citations
- Robust fine-tuning of zero-shot modelsMitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li et al.CVPR 2022 · 364 citations
- Hierarchical Decomposition of Prompt-Based Continual Learning: Rethinking Obscured Sub-optimalityLiyuan Wang, Jingyi Xie, Xingxing Zhang, Mingyi Huang et al.NeurIPS 2023 · 183 citations
Related papers
- Training Consistent Mixture-of-Experts-Based Prompt Generator for Continual LearningYue Lu, Shizhou Zhang, De Cheng, Guoqiang Liang et al.AAAI 2025 · 8 citations
- Theory on Mixture-of-Experts in Continual LearningHongbo Li, Sen Lin, Lingjie Duan, Yingbin Liang et al.ICLR 2025
- Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE AdaptationJunzhuo Li, Bo Wang, Xiuze Zhou, Xuming HuEMNLP 2025 · 5 citations
- Multi-Head Attention as a Source of Catastrophic Forgetting in MoE TransformersAnrui Chen, Ruijun Huang, Xin Zhang, Fang DONG(董方) et al.ICML 2026 · 3 citations
- One-Prompt Strikes Back: Sparse Mixture of Experts for Prompt-based Continual LearningMinh Le, Bao-Ngoc Dao, Huy Nguyen, Quyen Tran et al.ICLR 2026 · 3 citations
