Distillation-Guided Structural Transfer for Continual Learning Beyond Sparse Distributed Memory
Huiyan Xue, Xuming Ran, Yaxin Li, Qi Xu, Enhui Li, Yi Xu, Qiang Zhang
Abstract
Sparse neural systems are gaining traction for efficient continual learning due to their modularity and low interference. Architectures like Sparse Distributed Memory Multi-Layer Perceptrons (SDMLP) construct task-specific subnetworks via Top-K activation and have shown resilience against catastrophic forgetting. However, their rigid modularity poses two fundamental challenges: (1) the isolation of sparse subnetworks severely limits cross-task knowledge reuse; and (2) increased sparsity reduces interference but often degrades performance due to constrained feature sharing. We propose Selective Subnetwork Distillation (SSD), a structurally guided continual learning framework that treats distillation not as a regularizer, but as a topology-aligned information conduit. By identifying neurons with high activation frequency, SSD selectively distills knowledge within previous Top-K subnetworks and output logits-without requiring replay or task labels-preserving both sparsity and functional specialization. Unlike conventional distillation, SSD operates under hard modular constraints and enables structural realignment without altering the sparse architecture. While our method is validated on SDMLP, its structure-aligned mechanism has the potential to generalize to other sparse networks as a plug-in module for promoting representation sharing. Comprehensive experiments on Split CIFAR-10, CIFAR-100, and MNIST demonstrate that SSD improves accuracy, retention, and manifold coverage, offering a structurally grounded solution to sparse continual learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 50b670b9-cd87-4b1f-955c-adc818a50bf6Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du et al.NeurIPS 2022 · 933 citations
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
- Supermasks in SuperpositionMitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi et al.NeurIPS 2020 · 364 citations
- Training Neural Networks with Fixed Sparse MasksYi-Lin Sung, Varun Nair, Colin RaffelNeurIPS 2021 · 295 citations
Related papers
- Sparse Distributed Memory is a Continual LearnerTrenton Bricken, Xander Davies, Deepak Singh, Dmitry Krotov et al.ICLR 2023 · 5 citations
- Learning Bayesian Sparse Networks with Full Experience Replay for Continual LearningQingsen Yan, Dong Gong, Yuhang Liu, Anton van den Hengel et al.CVPR 2022 · 38 citations
- Dynamic Sub-graph Distillation for Robust Semi-supervised Continual LearningYan Fan, Yu Wang, Pengfei Zhu, Qinghua HuAAAI 2024 · 14 citations
- Remember the Past: Distilling Datasets into Addressable Memories for Neural NetworksZhiwei Deng, Olga RussakovskyNeurIPS 2022 · 140 citations
- Growing a Brain with Sparsity-Inducing Generation for Continual LearningHyundong Jin, Gyeong-Hyeon Kim, Chanho Ahn, Eunwoo KimICCV 2023 · 7 citations
