BECAME: Bayesian Continual Learning with Adaptive Model Merging
Mei Li, Yuxiang Lu, Qinyan Dai, Suizhi Huang, Yue Ding, Hongtao Lu
Abstract
Continual Learning (CL) strives to learn incrementally across tasks while mitigating catastrophic forgetting. A key challenge in CL is balancing stability (retaining prior knowledge) and plasticity (learning new tasks). While representative gradient projection methods ensure stability, they often limit plasticity. Model merging techniques offer promising solutions, but prior methods typically rely on empirical assumptions and carefully selected hyperparameters. In this paper, we explore the potential of model merging to enhance the stability-plasticity trade-off, providing theoretical insights that underscore its benefits. Specifically, we reformulate the merging mechanism using Bayesian continual learning principles and derive a closed-form solution for the optimal merging coefficient based on the Laplace approximation that adapts to the diverse characteristics of tasks. To validate our approach, we introduce a two-stage framework named BE-CAME, which synergizes the expertise of gradient projection and adaptive merging. Extensive experiments show that our approach outperforms state-of-the-art CL methods and existing merging strategies. Code is available at https: //github.com/limei0818/BECAME .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- TextResNet: Decoupling and Routing Optimization Signals in Compound AI Systems via Deep Residual TuningSuizhi Huang, Mei Li, Han Yu, Xiaoxiao LiICML 2026 · 3 citations
- EvoGM: Learning to Merge LLMs via Evolutionary Generative OptimizationTao Jiang, Xinmeng Yu, Chenhao Yi, Yiling Wu et al.ICML 2026 · 1 citation
- Beyond Buffer Limits: Energy-Based Data Reassembly for Continual LearningZhenyi Wang, Yixuan Sun, Yue Wang, Zhong Chen et al.ICML 2026
- Unlocking the Potential of Continual Model Merging: An ODE PerspectiveLihong Lin, Haidong KangICML 2026
- Task-Driven Subspace Decomposition for Knowledge Sharing and Isolation in LoRA-based Continual LearningLingfeng He, De Cheng, Huaijie Wang, Xi Yang et al.ICML 2026
Builds on22
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel et al.NeurIPS 2023 · 999 citations
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 741 citations
- What is being transferred in transfer learning?Behnam Neyshabur, Hanie Sedghi, Chiyuan ZhangNeurIPS 2020 · 654 citations
- Gradient Projection Memory for Continual LearningGobinda Saha, Isha Garg, Kaushik RoyICLR 2021 · 409 citations
Related papers
- Train with Perturbation, Infer after Merging: A Two-Stage Framework for Continual LearningHaomiao Qiu, Miao Zhang, Ziyue Qiao, Liqiang NieNeurIPS 2025 · 8 citations
- On the Stability-Plasticity Dilemma in Continual Meta-Learning: Theory and AlgorithmQi Chen, Changjian Shui, Ligong Han, Mario MarchandNeurIPS 2023 · 32 citations
- Recall-Oriented Continual Learning with Generative Adversarial Meta-ModelHaneol Kang, Dong-Wan ChoiAAAI 2024 · 3 citations
- Parameter Merging with Gradient-Guided Supermasks in Online Continual LearningBenliu Qiu, Heqian Qiu, Lanxiao Wang, Taijin Zhao et al.AAAI 2026
- Adapt Before Continual LearningAojun Lu, Tao Feng, Hangjie Yuan, Chunhui Ding et al.AAAI 2026
