Parameter Merging with Gradient-Guided Supermasks in Online Continual Learning
Benliu Qiu, Heqian Qiu, Lanxiao Wang, Taijin Zhao, Yu Dai, Lili Pan, Hongliang Li
Abstract
Online continual learning (OCL) aims at learning a non-stationary data stream in a way of reading each data sample only once, and hence suffers from the trade-off of catastrophic forgetting and insufficient learning. In this work, we firstly analytically establish relationship between loss functions and model parameters from the Bayesian perspective. Based on our analysis, we subsequently propose a parameter merging method with gradient-guided supermasks. Our method leverages 1-order and 2-order gradient information to construct supermasks that determine the merging weights between the old and new models. Our method performs direct arithmetic operations on parameters to update models, beyond traditional gradient descent. We further discover that a widely-used premise that 1-order gradients can be negligible is invalid in OCL, due to slow convergence incurred by insufficient learning. Additionally, we utilize a dual-model dual-view distillation strategy that can align output distributions of the new and merged models for each sample, further enhancing model performance. Extensive experiments are conducted on four benchmarks in OCL settings, including CIFAR-10, CIFAR-100, Tiny-ImageNet, and ImageNet-100. Experimental results demonstrate that our method is effective, and achieves a substantial boost over previous methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on23
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati et al.NeurIPS 2020 · 1,494 citations
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 741 citations
- Supermasks in SuperpositionMitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi et al.NeurIPS 2020 · 364 citations
- New Insights on Reducing Abrupt Representation Change in Online Continual LearningLucas Caccia, Rahaf Aljundi, Nader Asadi, Tinne Tuytelaars et al.ICLR 2022 · 279 citations
Related papers
- Non-exemplar Online Class-Incremental Continual Learning via Dual-Prototype Self-Augment and RefinementFushuo Huo, Wenchao Xu, Jingcai Guo, Haozhao Wang et al.AAAI 2024 · 25 citations
- BECAME: Bayesian Continual Learning with Adaptive Model MergingMei Li, Yuxiang Lu, Qinyan Dai, Suizhi Huang et al.ICML 2025
- Incremental Learning in Online ScenarioJiangpeng He, Runyu Mao, Zeman Shao, Fengqing ZhuCVPR 2020
- Orchestrate Latent Expertise: Advancing Online Continual Learning with Multi-Level Supervision and Reverse Self-DistillationHongwei Yan, Liyuan Wang, Kaisheng Ma, Yi ZhongCVPR 2024 · 15 citations
- Rethinking Momentum Knowledge Distillation in Online Continual LearningNicolas Michel, Maorong Wang, Ling Xiao, Toshihiko YamasakiICML 2024 · 26 citations
