VariGrow: Variational Architecture Growing for Task-Agnostic Continual Learning based on Bayesian Novelty
Randy Ardywibowo, Zepeng Huo, Zhangyang Wang, Bobak J. Mortazavi, Shuai Huang, Xiaoning Qian
Abstract
Continual Learning (CL) is the problem of sequentially learning predictive models with varying data that may originate from different contexts. Many existing CL methods assume that the data stream is divided into a sequence of contexts, termed as tasks, with explicitly given transition boundaries. Unfortunately, many real-world CL scenarios have neither explicit task information nor context boundaries, motivating the study of task-agnostic CL. This paper proposes a variational architecture growing framework dubbed VariGrow. By interpreting dynamically growing neural networks as a Bayesian approximation, and defining flexible implicit variational distributions, VariGrow detects if a new task is arriving through an energy-based novelty score. If the novelty score is high and the sample is "detected" as a new task, VariGrow will grow a new expert module to be responsible for it. Otherwise, the sample will be assigned to one of the existing experts who is the most "familiar" with it (i.e., one with the lowest novelty score) to preserve all the acquired knowledge. We have tested VariGrow on several CIFAR and ImageNet-based benchmarks for the strictly task-agnostic CL setting without any task information during training or testing, which demonstrates its consistently superior or competitive performance. More interesting, VariGrow achieves comparable performance with task-aware CL methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e4546b98-b091-47f8-8e97-7af7cc567892Cited by top-tier papers3
- Mitigating Catastrophic Forgetting in Online Continual Learning by Modeling Previous Task Interrelations via Pareto OptimizationYichen Wu, Hong Wang, Peilin Zhao, Yefeng Zheng et al.ICML 2024 · 23 citations
- Learning Expressive Priors for Generalization and Uncertainty Estimation in Neural NetworksDominik Schnaus, Jongseok Lee, Daniel Cremers, Rudolph TriebelICML 2023 · 5 citations
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic CorporaTuan-Luc Huynh, Thuy-Trang Vu, Weiqing Wang, Trung Le et al.EMNLP 2025 · 1 citation
Builds on13
- Energy-based Out-of-distribution DetectionWeitang Liu, Xiaoyun Wang, John D. Owens, Yixuan LiNeurIPS 2020 · 2,213 citations
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud et al.ICLR 2020 · 643 citations
- Supermasks in SuperpositionMitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi et al.NeurIPS 2020 · 364 citations
- A Neural Dirichlet Process Mixture Model for Task-Free Continual LearningSoochan Lee, Junsoo Ha, Dongsu Zhang, Gunhee KimICLR 2020 · 238 citations
- Self-Supervised Learning for Generalizable Out-of-Distribution DetectionSina Mohseni, Mandar Pitale, J. B. S. Yadawa, Zhangyang WangAAAI 2020 · 229 citations
Related papers
- Bayesian Structural Adaptation for Continual LearningAbhishek Kumar, Sunabha Chatterjee, Piyush RaiICML 2021 · 7 citations
- Self-Evolved Dynamic Expansion Model for Task-Free Continual LearningFei Ye, Adrian G. BorsICCV 2023 · 28 citations
- Continual Learning with Adaptive Weights (CLAW)Tameem Adel, Han Zhao, Richard E. TurnerICLR 2020 · 79 citations
- Continual Learning via Sequential Function-Space Variational InferenceTim G. J. Rudner, Freddie Bickford Smith, Qixuan Feng, Yee Whye Teh et al.ICML 2022 · 57 citations
- Wasserstein Expansible Variational Autoencoder for Discriminative and Generative Continual LearningFei Ye, Adrian G. BorsICCV 2023 · 6 citations
