Continual Learning With Extended Kronecker-Factored Approximate Curvature
Janghyeon Lee, Hyeong Gwon Hong, Donggyu Joo, Junmo Kim
Abstract
We propose a quadratic penalty method for continual learning of neural networks that contain batch normalization (BN) layers. The Hessian of a loss function represents the curvature of the quadratic penalty function, and a Kronecker-factored approximate curvature (K-FAC) is used widely to practically compute the Hessian of a neural network. However, the approximation is not valid if there is dependence between examples, typically caused by BN layers in deep network architectures. We extend the K-FAC method so that the inter-example relations are taken into account and the Hessian of deep neural networks can be properly approximated under practical assumptions. We also propose a method of weight merging and reparameterization to properly handle statistical parameters of BN, which plays a critical role for continual learning with BN, and a method that selects hyperparameters without source task data. Our method shows better performance than baselines in the permuted MNIST task with BN layers and in sequential learning from the ImageNet classification task to fine-grained classification tasks with ResNet-50, without any explicit or implicit use of source task data for hyperparameter selection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6da315b9-8d5b-4b22-bee7-90cb8c11b025Cited by top-tier papers15
- Tangent Model Composition for Ensembling and Continual Fine-tuningTian Yu Liu, Stefano SoattoICCV 2023 · 29 citations
- Detect, Distill and Update: Learned DB Systems Facing Out of Distribution DataMeghdad Kurmanji, Peter TriantafillouSIGMOD 2023 · 19 citations
- Interactive Continual Learning: Fast and Slow ThinkingBiqing Qi, Xinquan Chen, Junqi Gao, Dong Li et al.CVPR 2024 · 15 citations
- DACAPO: Accelerating Continuous Learning in Autonomous Systems for Video AnalyticsYoonsung Kim, Changhun Oh, Jinwoo Hwang, Wonung Kim et al.ISCA 2024 · 13 citations
- Integrating Task-Specific and Universal Adapters for Pre-Trained Model-Based Class-Incremental LearningYan Wang, Da-Wei Zhou, Han-Jia YeICCV 2025 · 7 citations
Builds on1
Related papers
- Overcoming Recency Bias of Normalization Statistics in Continual Learning: Balance and AdaptationYilin Lyu, Liyuan Wang, Xingxing Zhang, Zicheng Sun et al.NeurIPS 2023 · 17 citations
- Practical Quasi-Newton Methods for Training Deep Neural NetworksDonald Goldfarb, Yi Ren, Achraf BahamouNeurIPS 2020 · 130 citations
- Residual Continual LearningJanghyeon Lee, Donggyu Joo, Hyeong Gwon Hong, Junmo KimAAAI 2020 · 25 citations
- Convolutional neural network training with distributed K-FACJ. Gregory Pauloski, Zhao Zhang, Lei Huang, Weijia Xu et al.SC 2020 · 26 citations
- SKFAC: Training Neural Networks With Faster Kronecker-Factored Approximate CurvatureZedong Tang, Fenlong Jiang, Maoguo Gong, Hao Li et al.CVPR 2021
