Overcoming Recency Bias of Normalization Statistics in Continual Learning: Balance and Adaptation
Yilin Lyu, Liyuan Wang, Xingxing Zhang, Zicheng Sun, Hang Su, Jun Zhu, Liping Jing
Abstract
Continual learning entails learning a sequence of tasks and balancing their knowledge appropriately. With limited access to old training samples, much of the current work in deep neural networks has focused on overcoming catastrophic forgetting of old tasks in gradient-based optimization. However, the normalization layers provide an exception, as they are updated interdependently by the gradient and statistics of currently observed training samples, which require specialized strategies to mitigate recency bias. In this work, we focus on the most popular Batch Normalization (BN) and provide an in-depth theoretical analysis of its sub-optimality in continual learning. Our analysis demonstrates the dilemma between balance and adaptation of BN statistics for incremental tasks, which potentially affects training stability and generalization. Targeting on these particular challenges, we propose Adaptive Balance of BN (AdaBN), which incorporates appropriately a Bayesian-based strategy to adapt task-wise contributions and a modified momentum to balance BN statistics, corresponding to the training and testing stages. By implementing BN in a continual learning fashion, our approach achieves significant performance gains across a wide range of benchmarks, particularly for the challenging yet realistic online scenarios (e.g., up to 7.68%, 6.86% and 4.26% on Split CIFAR-10, Split CIFAR-100 and Split Mini-ImageNet, respectively). Our code is available at https://github.com/lvyilin/AdaB2N.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Orchestrate Latent Expertise: Advancing Online Continual Learning with Multi-Level Supervision and Reverse Self-DistillationHongwei Yan, Liyuan Wang, Kaisheng Ma, Yi ZhongCVPR 2024 · 15 citations
- What Matters in Graph Class Incremental Learning? An Information Preservation PerspectiveJialu Li, Yu Wang, Pengfei Zhu, Wanyu Lin et al.NeurIPS 2024 · 14 citations
- Accelerating Heterogeneous Federated Learning with Closed-form ClassifiersEros Fanì, Raffaello Camoriano, Barbara Caputo, Marco CicconeICML 2024 · 10 citations
- Continual Unsupervised Generative Modelling via Online Optimal TransportFei Ye, Adrian G. Bors, Kun ZhangAAAI 2025
Builds on10
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang et al.ICCV 2019 · 2,239 citations
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati et al.NeurIPS 2020 · 1,494 citations
- Gradient Projection Memory for Continual LearningGobinda Saha, Isha Garg, Kaushik RoyICLR 2021 · 409 citations
- New Insights on Reducing Abrupt Representation Change in Online Continual LearningLucas Caccia, Rahaf Aljundi, Nader Asadi, Tinne Tuytelaars et al.ICLR 2022 · 279 citations
- Memory Replay with Data Compression for Continual LearningLiyuan Wang, Xingxing Zhang, Kuo Yang, Longhui Yu et al.ICLR 2022 · 136 citations
Related papers
- Continual Normalization: Rethinking Batch Normalization for Online Continual LearningQuang Pham, Chenghao Liu, Steven C. H. HoiICLR 2022 · 72 citations
- Rebalancing Batch Normalization for Exemplar-Based Class-Incremental LearningSungmin Cha, Sungjun Cho, Dasol Hwang, Sunwon Hong et al.CVPR 2023
- New Insights for the Stability-Plasticity Dilemma in Online Continual LearningDahuin Jung, Dongjin Lee, Sunwon Hong, Hyemi Jang et al.ICLR 2023 · 4 citations
- Layerwise Optimization by Gradient Decomposition for Continual LearningShixiang Tang, Dapeng Chen, Jinguo Zhu, Shijie Yu et al.CVPR 2021
- Mitigating Forgetting in Online Continual Learning with Neuron CalibrationHaiyan Yin, Peng Yang, Ping LiNeurIPS 2021 · 39 citations
