Continual Normalization: Rethinking Batch Normalization for Online Continual Learning
Quang Pham, Chenghao Liu, Steven C. H. Hoi
Abstract
Existing continual learning methods use Batch Normalization (BN) to facilitate training and improve generalization across tasks. However, the non-i.i.d and non-stationary nature of continual learning data, especially in the online setting, amplify the discrepancy between training and testing in BN and hinder the performance of older tasks. In this work, we study the cross-task normalization effect of BN in online continual learning where BN normalizes the testing data using moments biased towards the current task, resulting in higher catastrophic forgetting. This limitation motivates us to propose a simple yet effective method that we call Continual Normalization (CN) to facilitate training similar to BN while mitigating its negative effect. Extensive experiments on different continual learning algorithms and online scenarios show that CN is a direct replacement for BN and can provide substantial performance improvements. Our implementation is available at https://github.com/phquang/Continual-Normalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b98d06fd-ee32-4812-ac0b-135c25007c4fCited by top-tier papers15
- CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and SimplicityAditya Bhatt, Daniel Palenicek, Boris Belousov, Max Argus et al.ICLR 2024 · 106 citations
- MOS: Model Surgery for Pre-Trained Model-Based Class-Incremental LearningHai-Long Sun, Da-Wei Zhou, Hanbin Zhao, Le Gan et al.AAAI 2025 · 31 citations
- Summarizing Stream Data for Memory-Constrained Online Continual LearningJianyang Gu, Kai Wang, Wei Jiang, Yang YouAAAI 2024 · 30 citations
- Does Continual Learning Equally Forget All Parameters?Haiyan Zhao, Tianyi Zhou, Guodong Long, Jing Jiang et al.ICML 2023 · 21 citations
- Overcoming Recency Bias of Normalization Statistics in Continual Learning: Balance and AdaptationYilin Lyu, Liyuan Wang, Xingxing Zhang, Zicheng Sun et al.NeurIPS 2023 · 17 citations
Builds on7
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati et al.NeurIPS 2020 · 1,494 citations
- Continual learning with hypernetworksJohannes von Oswald, Christian Henning, João Sacramento, Benjamin F. GreweICLR 2020 · 412 citations
- Anatomy of Catastrophic Forgetting: Hidden Representations and Task SemanticsVinay Venkatesh Ramasesh, Ethan Dyer, Maithra RaghuICLR 2021 · 207 citations
- DualNet: Continual Learning, Fast and SlowQuang Pham, Chenghao Liu, Steven C. H. HoiNeurIPS 2021 · 192 citations
- TaskNorm: Rethinking Batch Normalization for Meta-LearningJohn Bronskill, Jonathan Gordon, James Requeima, Sebastian Nowozin et al.ICML 2020 · 93 citations
Related papers
- Rebalancing Batch Normalization for Exemplar-Based Class-Incremental LearningSungmin Cha, Sungjun Cho, Dasol Hwang, Sunwon Hong et al.CVPR 2023
- Online Continual Learning through Mutual Information MaximizationYiduo Guo, Bing Liu, Dongyan ZhaoICML 2022 · 139 citations
- Residual Continual LearningJanghyeon Lee, Donggyu Joo, Hyeong Gwon Hong, Junmo KimAAAI 2020 · 25 citations
- CBA: Improving Online Continual Learning via Continual Bias AdaptorQuanziang Wang, Renzhen Wang, Yichen Wu, Xixi Jia et al.ICCV 2023 · 25 citations
- Mitigating Forgetting in Online Continual Learning with Neuron CalibrationHaiyan Yin, Peng Yang, Ping LiNeurIPS 2021 · 39 citations
