KSM: Fast Multiple Task Adaption via Kernel-Wise Soft Mask Learning
Li Yang, Zhezhi He, Junshan Zhang, Deliang Fan
Abstract
Deep Neural Networks (DNN) could forget the knowledge about earlier tasks when learning new tasks, and this is known as catastrophic forgetting. While recent continual learning methods are capable of alleviating the catastrophic problem on toy-sized datasets, some issues still remain to be tackled when applying them in real-world problems. Recently, the fast mask-based learning method (e.g. piggyback (Mallya, Davis, and Lazebnik 2018)) is proposed to address these issues by learning only a binary element-wise mask in a fast manner, while keeping the backbone model fixed. However, the binary mask has limited modeling capacity for new tasks. A more recent work (Hung et al. 2019) proposes a compressgrow-based method (CPG) to achieve better accuracy for new tasks by partially training backbone model, but with orderhigher training cost, which makes it infeasible to be deployed into popular state-of-the-art edge-/mobile-learning. The primary goal of this work is to simultaneously achieve fast and high-accuracy multi task adaption in continual learning setting. Thus motivated, we propose a new training method called kernel-wise Soft Mask (KSM), which learns a kernelwise hybrid binary and real-value soft mask for each task, while using the same backbone model. Such a soft mask can be viewed as a superposition of a binary mask and a properly scaled real-value tensor, which offers a richer representation capability without low-level kernel support to meet the objective of low hardware overhead. We validate KSM on multiple benchmark datasets against recent state-of-the-art methods (e.g. Piggyback, Packnet, CPG, etc.), which shows good improvement in both accuracy and training cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- XMA: a crossbar-aware multi-task adaption framework via shift-based mask learning methodFan Zhang, Li Yang, Jian Meng, Jae-sun Seo et al.DAC 2022 · 3 citations
- Parameter-Level Soft-Masking for Continual LearningTatsuya Konishi, Mori Kurokawa, Chihiro Ono, Zixuan Ke et al.ICML 2023 · 63 citations
- FedKNOW: Federated Continual Learning with Signature Task Knowledge Integration at EdgeYaxin Luopan, Rui Han, Qinglong Zhang, Chi Harold Liu et al.ICDE 2023 · 31 citations
- Meta-attention for ViT-backed Continual LearningMengqi Xue, Haofei Zhang, Jie Song, Mingli SongCVPR 2022 · 40 citations
- SparCL: Sparse Continual Learning on the EdgeZifeng Wang, Zheng Zhan, Yifan Gong, Geng Yuan et al.NeurIPS 2022 · 97 citations
