Theory on Mixture-of-Experts in Continual Learning
Hongbo Li, Sen Lin, Lingjie Duan, Yingbin Liang, Ness B. Shroff
Abstract
Continual learning (CL) has garnered significant attention because of its ability to adapt to new tasks that arrive over time. Catastrophic forgetting (of old tasks) has been identified as a major issue in CL, as the model adapts to new tasks. The Mixture-of-Experts (MoE) model has recently been shown to effectively mitigate catastrophic forgetting in CL, by employing a gating network to sparsify and distribute diverse tasks among multiple experts. However, there is a lack of theoretical analysis of MoE and its impact on the learning performance in CL. This paper provides the first theoretical results to characterize the impact of MoE in CL via the lens of overparameterized linear regression tasks. We establish the benefit of MoE over a single expert by proving that the MoE model can diversify its experts to specialize in different tasks, while its router learns to select the right expert for each task and balance the loads across all experts. Our study further suggests an intriguing fact that the MoE in CL needs to terminate the update of the gating network after sufficient training rounds to attain system convergence, which is not needed in the existing MoE studies that do not consider the continual task arrival. Furthermore, we provide explicit expressions for the expected forgetting and overall generalization error to characterize the benefit of MoE in the learning performance in CL. Interestingly, adding more experts requires additional rounds before convergence, which may not enhance the learning performance. Finally, we conduct experiments on both synthetic and real datasets to extend these insights from linear models to deep neural networks (DNNs), which also shed light on the practical algorithm design for MoE in CL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ff1a1891-5682-4248-9e8b-63daa2e648fbCited by top-tier papers23
- Theory of Mixture-of-Experts for Mobile Edge ComputingHongbo Li, Lingjie DuanINFOCOM 2025 · 9 citations
- Multi-Head Attention as a Source of Catastrophic Forgetting in MoE TransformersAnrui Chen, Ruijun Huang, Xin Zhang, Fang DONG(董方) et al.ICML 2026 · 3 citations
- On the Expressive Power of Mixture-of-Experts for Structured Complex TasksMingze Wang, Weinan ENeurIPS 2025 · 3 citations
- On the Theory of Continual Learning with Gradient Descent for Neural NetworksHossein Taheri, Avishek Ghosh, Arya MazumdarICML 2026 · 2 citations
- Little By Little: Continual Learning via Incremental Mixture of Rank-1 Associative Memory ExpertsHaodong Lu, Chongyang Zhao, Minhui Xue, Lina Yao et al.ICML 2026 · 2 citations
Related papers
- Towards Understanding the Mixture-of-Experts Layer in Deep LearningZixiang Chen, Yihe Deng, Yue Wu, Quanquan Gu et al.NeurIPS 2022 · 199 citations
- Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE AdaptationJunzhuo Li, Bo Wang, Xiuze Zhou, Xuming HuEMNLP 2025 · 5 citations
- Mixing Expertise with Confidence: A Mixture of Experts Framework for Robust Multi-Modal Continual LearningMd Abdullah Al Forhad, Yuansheng Zhu Zhu., Abhinab Acharya, Xumin Liu et al.ICML 2026
- Spectral Mixture-of-Experts for Continual LearningChen Yin, Xingbo Dong, Xuelin Shen, Zhe JinCVPR 2026
- MINGLE: Mixture of Null-Space Gated Low-Rank Experts for Test-Time Continual Model MergingZihuan Qiu, Yi Xu, Chiyuan He, Fanman Meng et al.NeurIPS 2025 · 16 citations
