TinySubNets: An Efficient and Low Capacity Continual Learning Strategy
Marcin Pietron, Kamil Faber, Dominik Zurek, Roberto Corizzo
摘要
Continual Learning (CL) is a highly relevant setting gaining traction in recent machine learning research. Among CL works, architectural and hybrid strategies are particularly effective due to their potential to adapt the model architecture as new tasks are presented. However, many existing solutions do not efficiently exploit model sparsity, and are prone to capacity saturation due to their inefficient use of available weights, which limits the number of learnable tasks. In this paper, we propose TinySubNets (TSN), a novel architectural CL strategy that addresses the issues through the unique combination of pruning with different sparsity levels, adaptive quantization, and weight sharing. Pruning identifies a subset of weights that preserve model performance, making less relevant weights available for future tasks. Adaptive quantization allows a single weight to be separated into multiple parts which can be assigned to different tasks. Weight sharing between tasks boosts the exploitation of capacity and task similarity, allowing for the identification of a better trade-off between model accuracy and capacity. These features allow TSN to efficiently leverage the available capacity, enhance knowledge transfer, and reduce computational resources consumption. Experimental results involving common benchmark CL datasets and scenarios show that our proposed strategy achieves better results in terms of accuracy than existing state-of-the-art CL strategies. Moreover, our strategy is shown to provide a significantly improved model capacity exploitation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SAMCL: Empowering SAM to Continually Learn from Dynamic Domains with Extreme Storage EfficiencyZeqing Wang, Kangye Ji, Di Wang, Haibin Zhang 等AAAI 2026 · 被引用 2 次
- MaRS: Memory-Adaptive Routing for Reliable Capacity Expansion and Knowledge RetentionGang YanICLR 2026
它引用的顶会 Paper10
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati 等NeurIPS 2020 · 被引用 1,494 次
- Gradient Projection Memory for Continual LearningGobinda Saha, Isha Garg, Kaushik RoyICLR 2021 · 被引用 409 次
- Supermasks in SuperpositionMitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi 等NeurIPS 2020 · 被引用 364 次
- DyTox: Transformers for Continual Learning with DYnamic TOken eXpansionArthur Douillard, Alexandre Ramé, Guillaume Couairon, Matthieu CordCVPR 2022 · 被引用 315 次
- RanPAC: Random Projections and Pre-trained Models for Continual LearningMark D. McDonnell, Dong Gong, Amin Parvaneh, Ehsan Abbasnejad 等NeurIPS 2023 · 被引用 245 次
相关 Paper
- Task-aware Orthogonal Sparse Network for Exploring Shared Knowledge in Continual LearningYusong Hu, De Cheng, Dingwen Zhang, Nannan Wang 等ICML 2024 · 被引用 14 次
- Continual Learning with Adaptive Weights (CLAW)Tameem Adel, Han Zhao, Richard E. TurnerICLR 2020 · 被引用 79 次
- Parameter-Level Soft-Masking for Continual LearningTatsuya Konishi, Mori Kurokawa, Chihiro Ono, Zixuan Ke 等ICML 2023 · 被引用 63 次
- Learning Bayesian Sparse Networks with Full Experience Replay for Continual LearningQingsen Yan, Dong Gong, Yuhang Liu, Anton van den Hengel 等CVPR 2022 · 被引用 38 次
- Forget-free Continual Learning with Winning SubnetworksHaeyong Kang, Rusty John Lloyd Mina, Sultan Rizky Hikmawan Madjid, Jaehong Yoon 等ICML 2022 · 被引用 159 次
