Heterogeneous Continual Learning
Divyam Madaan, Hongxu Yin, Wonmin Byeon, Jan Kautz, Pavlo Molchanov
Abstract
We propose a novel framework and a solution to tackle the continual learning (CL) problem with changing network architectures. Most CL methods focus on adapting a single architecture to a new task/class by modifying its weights. However, with rapid progress in architecture design, the problem of adapting existing solutions to novel architectures becomes relevant. To address this limitation, we propose Heterogeneous Continual Learning (HCL), where a wide range of evolving network architectures emerge continually together with novel data/tasks. As a solution, we build on top of the distillation family of techniques and modify it to a new setting where a weaker model takes the role of a teacher; meanwhile, a new stronger architecture acts as a student. Furthermore, we consider a setup of limited access to previous data and propose Quick Deep Inversion (QDI) to recover prior task visual features to support knowledge transfer. QDI significantly reduces computational costs compared to previous solutions and improves overall performance. In summary, we propose a new setup for CL with a modified knowledge distillation paradigm and design a quick data inversion method to enhance distillation. Our evaluation of various benchmarks shows a significant improvement on accuracy in comparison to state-of-the-art methods over various networks architectures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- TinySubNets: An Efficient and Low Capacity Continual Learning StrategyMarcin Pietron, Kamil Faber, Dominik Zurek, Roberto CorizzoAAAI 2025 · 6 citations
- Reawakening knowledge: Anticipatory recovery from catastrophic interference via structured trainingYanlai Yang, Matt Jones, Michael C. Mozer, Mengye RenNeurIPS 2024 · 6 citations
Builds on15
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati et al.NeurIPS 2020 · 1,494 citations
Related papers
- Continual Learning with Adaptive Weights (CLAW)Tameem Adel, Han Zhao, Richard E. TurnerICLR 2020 · 79 citations
- Split-and-Bridge: Adaptable Class Incremental Learning within a Single Neural NetworkJong-Yeong Kim, Dong-Wan ChoiAAAI 2021 · 28 citations
- Remember the Past: Distilling Datasets into Addressable Memories for Neural NetworksZhiwei Deng, Olga RussakovskyNeurIPS 2022 · 140 citations
- Class Similarity Weighted Knowledge Distillation for Continual Semantic SegmentationMinh-Hieu Phan, The-Anh Ta, Son Lam Phung, Long Tran-Thanh et al.CVPR 2022 · 57 citations
- On Generalizing Beyond Domains in Cross-Domain Continual LearningChristian Simon, Masoud Faraki, Yi-Hsuan Tsai, Xiang Yu et al.CVPR 2022 · 34 citations
