Mixture of Experts Meets Prompt-Based Continual Learning
Minh Le, An Nguyen The, Huy Nguyen, Trang Nguyen, Trang Pham, Linh Ngo Van, Nhat Ho
摘要
Exploiting the power of pre-trained models, prompt-based approaches stand out compared to other continual learning solutions in effectively preventing catastrophic forgetting, even with very few learnable parameters and without the need for a memory buffer. While existing prompt-based continual learning methods excel in leveraging prompts for state-of-the-art performance, they often lack a theoretical explanation for the effectiveness of prompting. This paper conducts a theoretical analysis to unravel how prompts bestow such advantages in continual learning, thus offering a new perspective on prompt design. We first show that the attention block of pre-trained models like Vision Transformers inherently encodes a special mixture of experts architecture, characterized by linear experts and quadratic gating score functions. This realization drives us to provide a novel view on prefix tuning, reframing it as the addition of new task-specific experts, thereby inspiring the design of a novel gating mechanism termed Non-linear Residual Gates (NoRGa). Through the incorporation of non-linear activation and residual connection, NoRGa enhances continual learning performance while preserving parameter efficiency. The effectiveness of NoRGa is substantiated both theoretically and empirically across diverse benchmarks and pretraining paradigms. Our code is publicly available at https://github.com/Minhchuyentoancbn/MoE_PromptCL
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- Gated Integration of Low-Rank Adaptation for Continual Learning of Large Language ModelsYan-Shuo Liang, Jia-Rui Chen, Wu-Jun LiNeurIPS 2025 · 被引用 15 次
- SEMPO: Lightweight Foundation Models for Time Series ForecastingHui He, Kun Yi, Yuanchi Ma, Qi Zhang 等NeurIPS 2025 · 被引用 12 次
- Revisit Visual Prompt Tuning: The Expressiveness of Prompt ExpertsMinh Le, Anh Nguyen, Huy Nguyen, Chau Nguyen 等ICLR 2026 · 被引用 6 次
- FlyPrompt: Brain-Inspired Random-Expanded Routing with Temporal-Ensemble Experts for General Continual LearningHongwei Yan, Guanglong Sun, Kanglei Zhou, Qian Li 等ICLR 2026 · 被引用 4 次
- Forging a Dynamic Memory: Retrieval-Guided Continual Learning for Generalist Medical Foundation ModelsZizhi Chen, Yizhen Gao, Minghao Han, Yizhou Liu 等CVPR 2026 · 被引用 3 次
它引用的顶会 Paper17
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 被引用 2,340 次
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang 等CVPR 2022 · 被引用 635 次
- S-Prompts Learning with Pre-trained Transformers: An Occam's Razor for Domain Incremental LearningYabin Wang, Zhiwu Huang, Xiaopeng HongNeurIPS 2022 · 被引用 397 次
- Co2L: Contrastive Continual LearningHyuntak Cha, Jaeho Lee, Jinwoo ShinICCV 2021 · 被引用 391 次
相关 Paper
- Visual Prompt Tuning in Null Space for Continual LearningYue Lu, Shizhou Zhang, De Cheng, Yinghui Xing 等NeurIPS 2024 · 被引用 42 次
- CODA-Prompt: COntinual Decomposed Attention-Based Prompting for Rehearsal-Free Continual LearningJames Seale Smith, Leonid Karlinsky, Vyshnavi Gutta, Paola Cascante-Bonilla 等CVPR 2023
- Convolutional Prompting meets Language Models for Continual LearningAnurag Roy, Riddhiman Moulick, Vinay Kumar Verma, Saptarshi Ghosh 等CVPR 2024 · 被引用 15 次
- Prompt Gradient Projection for Continual LearningJingyang Qiao, Zhizhong Zhang, Xin Tan, Chengwei Chen 等ICLR 2024 · 被引用 47 次
- Self-Regulating Prompt Expansion for Continual LearningYiwen Wang, Diana Benavides-Prado, Yun Sing KohKDD 2026
