BECAME: Bayesian Continual Learning with Adaptive Model Merging
Mei Li, Yuxiang Lu, Qinyan Dai, Suizhi Huang, Yue Ding, Hongtao Lu
摘要
Continual Learning (CL) strives to learn incrementally across tasks while mitigating catastrophic forgetting. A key challenge in CL is balancing stability (retaining prior knowledge) and plasticity (learning new tasks). While representative gradient projection methods ensure stability, they often limit plasticity. Model merging techniques offer promising solutions, but prior methods typically rely on empirical assumptions and carefully selected hyperparameters. In this paper, we explore the potential of model merging to enhance the stability-plasticity trade-off, providing theoretical insights that underscore its benefits. Specifically, we reformulate the merging mechanism using Bayesian continual learning principles and derive a closed-form solution for the optimal merging coefficient based on the Laplace approximation that adapts to the diverse characteristics of tasks. To validate our approach, we introduce a two-stage framework named BE-CAME, which synergizes the expertise of gradient projection and adaptive merging. Extensive experiments show that our approach outperforms state-of-the-art CL methods and existing merging strategies. Code is available at https: //github.com/limei0818/BECAME .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- TextResNet: Decoupling and Routing Optimization Signals in Compound AI Systems via Deep Residual TuningSuizhi Huang, Mei Li, Han Yu, Xiaoxiao LiICML 2026 · 被引用 3 次
- EvoGM: Learning to Merge LLMs via Evolutionary Generative OptimizationTao Jiang, Xinmeng Yu, Chenhao Yi, Yiling Wu 等ICML 2026 · 被引用 1 次
- Beyond Buffer Limits: Energy-Based Data Reassembly for Continual LearningZhenyi Wang, Yixuan Sun, Yue Wang, Zhong Chen 等ICML 2026
- Unlocking the Potential of Continual Model Merging: An ODE PerspectiveLihong Lin, Haidong KangICML 2026
- Task-Driven Subspace Decomposition for Knowledge Sharing and Isolation in LoRA-based Continual LearningLingfeng He, De Cheng, Huaijie Wang, Xi Yang 等ICML 2026
它引用的顶会 Paper22
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel 等NeurIPS 2023 · 被引用 999 次
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 被引用 741 次
- What is being transferred in transfer learning?Behnam Neyshabur, Hanie Sedghi, Chiyuan ZhangNeurIPS 2020 · 被引用 654 次
- Gradient Projection Memory for Continual LearningGobinda Saha, Isha Garg, Kaushik RoyICLR 2021 · 被引用 409 次
相关 Paper
- Train with Perturbation, Infer after Merging: A Two-Stage Framework for Continual LearningHaomiao Qiu, Miao Zhang, Ziyue Qiao, Liqiang NieNeurIPS 2025 · 被引用 8 次
- On the Stability-Plasticity Dilemma in Continual Meta-Learning: Theory and AlgorithmQi Chen, Changjian Shui, Ligong Han, Mario MarchandNeurIPS 2023 · 被引用 32 次
- Recall-Oriented Continual Learning with Generative Adversarial Meta-ModelHaneol Kang, Dong-Wan ChoiAAAI 2024 · 被引用 3 次
- Parameter Merging with Gradient-Guided Supermasks in Online Continual LearningBenliu Qiu, Heqian Qiu, Lanxiao Wang, Taijin Zhao 等AAAI 2026
- Adapt Before Continual LearningAojun Lu, Tao Feng, Hangjie Yuan, Chunhui Ding 等AAAI 2026
