Training Consistent Mixture-of-Experts-Based Prompt Generator for Continual Learning
Yue Lu, Shizhou Zhang, De Cheng, Guoqiang Liang, Yinghui Xing, Nannan Wang, Yanning Zhang
Abstract
Visual prompt tuning-based continual learning (CL) methods have shown promising performance in exemplar-free scenarios, where their key component can be viewed as a prompt generator. Existing approaches generally rely on freezing old prompts, slow updating and task discrimination for prompt generators to preserve stability and minimize forgetting. In contrast, we introduce a novel approach that trains a consistent prompt generator to ensure stability during CL. Consistency means that for any instance from an old task, its corresponding instance-ware prompt generated by the prompt generator remains consistent even as the generator continually updates in a new task. This ensures that the representation of a specific instance remains stable across tasks and thereby prevents forgetting. We employ a mixture of experts (MoE) as the prompt generator, which contains a router and multiple experts. By deriving conditions sufficient to achieve the consistency for the MoE prompt generator, we demonstrate that: during training in a new task, if the router and experts update in the directions orthogonal to the subspaces spanned by old input features and gating vectors, respectively, the consistency can be theoretically guaranteed. To implement this orthogonality, we project parameter gradients to those orthogonal directions using the orthogonal projection matrices computed via the null space method. Extensive experiments on four class-incremental learning benchmarks validate the effectiveness and superiority of our approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d8581bcc-3cd0-4ace-9b3a-e2f3ff1e8d57Cited by top-tier papers6
- Is Parameter Isolation Better for Prompt-Based Continual Learning?Jiangyang Li, Chenhao Ding, SongLin Dong, Qiang Wang et al.CVPR 2026
- Attention Retention for Continual Learning with Vision TransformersYue Lu, Xiangyu Zhou, Shizhou Zhang, Yinghui Xing et al.AAAI 2026
- SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction TuningZhen-Hao Xie Xie, Jun-Tao Tang, Yu-Cheng Shi, Han-Jia Ye et al.ICML 2026
- Spectral Mixture-of-Experts for Continual LearningChen Yin, Xingbo Dong, Xuelin Shen, Zhe JinCVPR 2026
- Task-Driven Subspace Decomposition for Knowledge Sharing and Isolation in LoRA-based Continual LearningLingfeng He, De Cheng, Huaijie Wang, Xi Yang et al.ICML 2026
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang et al.ICCV 2019 · 2,239 citations
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang et al.CVPR 2022 · 635 citations
Related papers
- Visual Prompt Tuning in Null Space for Continual LearningYue Lu, Shizhou Zhang, De Cheng, Yinghui Xing et al.NeurIPS 2024 · 42 citations
- Prompt Gradient Projection for Continual LearningJingyang Qiao, Zhizhong Zhang, Xin Tan, Chengwei Chen et al.ICLR 2024 · 47 citations
- Consistent Prompting for Rehearsal-Free Continual LearningZhanxin Gao, Jun Cen, Xiaobin ChangCVPR 2024
- KSS-MoE: Knowledge Space Synergy Framework in Mixture of Experts for Continual Visual Instruction TuningLingyun Song, Ziyao Chen, Kang Pan, Xiaolin Han et al.AAAI 2026
- Adaptive Prompting for Continual Relation Extraction: A Within-Task Variance PerspectiveMinh Le, Tien Ngoc Luu, An Nguyen The, Thanh-Thien Le et al.AAAI 2025 · 12 citations
