Scalable and Order-robust Continual Learning with Additive Parameter Decomposition
Jaehong Yoon, Saehoon Kim, Eunho Yang, Sung Ju Hwang
Abstract
While recent continual learning methods largely alleviate the catastrophic problem on toy-sized datasets, some issues remain to be tackled to apply them to real-world problem domains. First, a continual learning model should effectively handle catastrophic forgetting and be efficient to train even with a large number of tasks. Secondly, it needs to tackle the problem of order-sensitivity, where the performance of the tasks largely varies based on the order of the task arrival sequence, as it may cause serious problems where fairness plays a critical role (e.g. medical diagnosis). To tackle these practical challenges, we propose a novel continual learning method that is scalable as well as order-robust, which instead of learning a completely shared set of weights, represents the parameters for each task as a sum of task-shared and sparse task-adaptive parameters. With our Additive Parameter Decomposition (APD), the task-adaptive parameters for earlier tasks remain mostly unaffected, where we update them only to reflect the changes made to the task-shared parameters. This decomposition of parameters effectively prevents catastrophic forgetting and order-sensitivity, while being computation- and memory-efficient. Further, we can achieve even better scalability with APD using hierarchical knowledge consolidation, which clusters the task-adaptive parameters to obtain hierarchically shared parameters. We validate our network with APD, APD-Net, on multiple benchmark datasets against state-of-the-art continual learning methods, which it largely outperforms in accuracy, scalability, and order-robustness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1ab70d07-58d0-4646-ad23-6a94341f79fbCited by top-tier papers64
- Gradient Projection Memory for Continual LearningGobinda Saha, Isha Garg, Kaushik RoyICLR 2021 · 409 citations
- Federated Continual Learning with Weighted Inter-client TransferJaehong Yoon, Wonyong Jeong, Giwoong Lee, Eunho Yang et al.ICML 2021 · 303 citations
- Federated Semi-Supervised Learning with Inter-Client Consistency & Disjoint LearningWonyong Jeong, Jaehong Yoon, Eunho Yang, Sung Ju HwangICLR 2021 · 271 citations
- Online Coreset Selection for Rehearsal-based Continual LearningJaehong Yoon, Divyam Madaan, Eunho Yang, Sung Ju HwangICLR 2022 · 181 citations
- Forget-free Continual Learning with Winning SubnetworksHaeyong Kang, Rusty John Lloyd Mina, Sultan Rizky Hikmawan Madjid, Jaehong Yoon et al.ICML 2022 · 159 citations
Related papers
- Layerwise Optimization by Gradient Decomposition for Continual LearningShixiang Tang, Dapeng Chen, Jinguo Zhu, Shijie Yu et al.CVPR 2021
- Parameter-Level Soft-Masking for Continual LearningTatsuya Konishi, Mori Kurokawa, Chihiro Ono, Zixuan Ke et al.ICML 2023 · 63 citations
- Residual Continual LearningJanghyeon Lee, Donggyu Joo, Hyeong Gwon Hong, Junmo KimAAAI 2020 · 25 citations
- Memory Efficient Continual Learning with TransformersBeyza Ermis, Giovanni Zappella, Martin Wistuba, Aditya Rawal et al.NeurIPS 2022 · 75 citations
- Continual Learning with Adaptive Weights (CLAW)Tameem Adel, Han Zhao, Richard E. TurnerICLR 2020 · 79 citations
