Continual Learning With Extended Kronecker-Factored Approximate Curvature
Janghyeon Lee, Hyeong Gwon Hong, Donggyu Joo, Junmo Kim
摘要
We propose a quadratic penalty method for continual learning of neural networks that contain batch normalization (BN) layers. The Hessian of a loss function represents the curvature of the quadratic penalty function, and a Kronecker-factored approximate curvature (K-FAC) is used widely to practically compute the Hessian of a neural network. However, the approximation is not valid if there is dependence between examples, typically caused by BN layers in deep network architectures. We extend the K-FAC method so that the inter-example relations are taken into account and the Hessian of deep neural networks can be properly approximated under practical assumptions. We also propose a method of weight merging and reparameterization to properly handle statistical parameters of BN, which plays a critical role for continual learning with BN, and a method that selects hyperparameters without source task data. Our method shows better performance than baselines in the permuted MNIST task with BN layers and in sequential learning from the ImageNet classification task to fine-grained classification tasks with ResNet-50, without any explicit or implicit use of source task data for hyperparameter selection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Tangent Model Composition for Ensembling and Continual Fine-tuningTian Yu Liu, Stefano SoattoICCV 2023 · 被引用 29 次
- Detect, Distill and Update: Learned DB Systems Facing Out of Distribution DataMeghdad Kurmanji, Peter TriantafillouSIGMOD 2023 · 被引用 19 次
- Interactive Continual Learning: Fast and Slow ThinkingBiqing Qi, Xinquan Chen, Junqi Gao, Dong Li 等CVPR 2024 · 被引用 15 次
- DACAPO: Accelerating Continuous Learning in Autonomous Systems for Video AnalyticsYoonsung Kim, Changhun Oh, Jinwoo Hwang, Wonung Kim 等ISCA 2024 · 被引用 13 次
- Integrating Task-Specific and Universal Adapters for Pre-Trained Model-Based Class-Incremental LearningYan Wang, Da-Wei Zhou, Han-Jia YeICCV 2025 · 被引用 7 次
它引用的顶会 Paper1
相关 Paper
- Overcoming Recency Bias of Normalization Statistics in Continual Learning: Balance and AdaptationYilin Lyu, Liyuan Wang, Xingxing Zhang, Zicheng Sun 等NeurIPS 2023 · 被引用 17 次
- Practical Quasi-Newton Methods for Training Deep Neural NetworksDonald Goldfarb, Yi Ren, Achraf BahamouNeurIPS 2020 · 被引用 130 次
- Residual Continual LearningJanghyeon Lee, Donggyu Joo, Hyeong Gwon Hong, Junmo KimAAAI 2020 · 被引用 25 次
- Convolutional neural network training with distributed K-FACJ. Gregory Pauloski, Zhao Zhang, Lei Huang, Weijia Xu 等SC 2020 · 被引用 26 次
- SKFAC: Training Neural Networks With Faster Kronecker-Factored Approximate CurvatureZedong Tang, Fenlong Jiang, Maoguo Gong, Hao Li 等CVPR 2021
