Parameter-Level Soft-Masking for Continual Learning
Tatsuya Konishi, Mori Kurokawa, Chihiro Ono, Zixuan Ke, Gyuhak Kim, Bing Liu
Abstract
Existing research on task incremental learning in continual learning has primarily focused on preventing catastrophic forgetting (CF). Although several techniques have achieved learning with no CF, they attain it by letting each task monopolize a sub-network in a shared network, which seriously limits knowledge transfer (KT) and causes over-consumption of the network capacity, i.e., as more tasks are learned, the performance deteriorates. The goal of this paper is threefold: (1) overcoming CF, (2) encouraging KT, and (3) tackling the capacity problem. A novel technique (called SPG) is proposed that soft-masks (partially blocks) parameter updating in training based on the importance of each parameter to old tasks. Each task still uses the full network, i.e., no monopoly of any part of the network by any task, which enables maximum KT and reduction in capacity usage. To our knowledge, this is the first work that soft-masks a model at the parameter-level for continual learning. Extensive experiments demonstrate the effectiveness of SPG in achieving all three objectives. More notably, it attains significant transfer of knowledge not only among similar tasks (with shared knowledge) but also among dissimilar tasks (with little shared knowledge) while mitigating CF.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c4312f5-bd99-4dcc-9f62-ae0f0f25053eCited by top-tier papers19
- Make Continual Learning Stronger via C-FlatAng Bian, Wei Li, Hangjie Yuan, Chengrong Yu et al.NeurIPS 2024 · 48 citations
- Vector Quantization Prompting for Continual LearningLi Jiao, Qiuxia Lai, Yu Li, Qiang XuNeurIPS 2024 · 15 citations
- Layerwise Proximal Replay: A Proximal Point Method for Online Continual LearningJinsoo Yoo, Yunpeng Liu, Frank Wood, Geoff PleissICML 2024 · 13 citations
- Prospective Representation Learning for Non-Exemplar Class-Incremental LearningWuxuan Shi, Mang YeNeurIPS 2024 · 9 citations
- Bayesian Adaptation of Network Depth and Width for Continual LearningJeevan Thapa, Rui LiICML 2024 · 7 citations
Builds on4
- Supermasks in SuperpositionMitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi et al.NeurIPS 2020 · 364 citations
- Scalable and Order-robust Continual Learning with Additive Parameter DecompositionJaehong Yoon, Saehoon Kim, Eunho Yang, Sung Ju HwangICLR 2020 · 206 citations
- Continual Learning of a Mixed Sequence of Similar and Dissimilar TasksZixuan Ke, Bing Liu, Xingchang HuangNeurIPS 2020 · 173 citations
- Forget-free Continual Learning with Winning SubnetworksHaeyong Kang, Rusty John Lloyd Mina, Sultan Rizky Hikmawan Madjid, Jaehong Yoon et al.ICML 2022 · 159 citations
Related papers
- BNS: Building Network Structures Dynamically for Continual LearningQi Qin, Wenpeng Hu, Han Peng, Dongyan Zhao et al.NeurIPS 2021 · 54 citations
- Task-aware Orthogonal Sparse Network for Exploring Shared Knowledge in Continual LearningYusong Hu, De Cheng, Dingwen Zhang, Nannan Wang et al.ICML 2024 · 14 citations
- Residual Continual LearningJanghyeon Lee, Donggyu Joo, Hyeong Gwon Hong, Junmo KimAAAI 2020 · 25 citations
- Achieving Forgetting Prevention and Knowledge Transfer in Continual LearningZixuan Ke, Bing Liu, Nianzu Ma, Hu Xu et al.NeurIPS 2021 · 167 citations
- Continual Learning with Node-Importance based Adaptive Group Sparse RegularizationSangwon Jung, Hongjoon Ahn, Sungmin Cha, Taesup MoonNeurIPS 2020 · 176 citations
