Efficient Bilevel Optimization for CKA-Guided MoE Upcycling
Zhiyuan Yu, Enneng Yang, Hao Jiang, Guojie Zhu, Feihong He, Peng Wang, Li Shen
Abstract
Upcycling, a strategy that initializes Mixture-of-Experts (MoE) by replicating pre-trained feed-forward or MoE networks to expand model capacity, has become a popular method in continual learning due to its effectiveness in mitigating catastrophic forgetting. However, existing paradigms indiscriminately expand capacity to prioritize performance at the cost of severe inefficiency, introducing severe parameter redundancy and failing to exploit structural heterogeneity. To address this, we investigate the determinants of forgetting in training dynamics using Centered Kernel Alignment (CKA) and loss landscape flatness to analyze the behavior of pre- and post-expansion MoE layers, uncovering instability in deep-layer representations and heterogeneous expert sensitivity to new tasks, thereby demonstrating the potential of selective upcycling to eliminate redundancy. Consequently, we propose a dynamic bilevel optimization framework to guide adaptive upcycling, featuring an outer loop employing a Gumbel-Softmax differentiable mask to perform Neural Architecture Search (NAS) for adaptive growth, while an inner loop optimizes weight updates via task objectives and CKA-regularized replay. Experiments on the TRACE benchmark demonstrate that our proposed method achieves better average accuracy with 80% forgetting reduction, while effectively eliminating 60% of redundant parameter expansion that standard upcycling would introduce.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Do Wide and Deep Networks Learn the Same Things? Uncovering How Neural Network Representations Vary with Width and DepthThao Nguyen, Maithra Raghu, Simon KornblithICLR 2021 · 323 citations
- Lifelong Language Pretraining with Distribution-Specialized ExpertsWuyang Chen, Yanqi Zhou, Nan Du, Yanping Huang et al.ICML 2023 · 85 citations
Related papers
- Online Continual Learning via Dynamic Expandable Recursive ModelFei Ye, Adrian G. BorsACM MM 2025
- Grow-on-Demand: Sparse and Adaptive Expert Expansion for Continual Instruction TuningYing Zhang, Xingyue Guo, Yu Zhao, Xuhui Sui et al.AAAI 2026
- Spectral Mixture-of-Experts for Continual LearningChen Yin, Xingbo Dong, Xuelin Shen, Zhe JinCVPR 2026
- Theory on Mixture-of-Experts in Continual LearningHongbo Li, Sen Lin, Lingjie Duan, Yingbin Liang et al.ICLR 2025
- Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE AdaptationJunzhuo Li, Bo Wang, Xiuze Zhou, Xuming HuEMNLP 2025 · 5 citations
