Continual learning with the neural tangent ensemble
Ari S. Benjamin, Christian-Gernot Pehle, Kyle Daruwalla
Abstract
A natural strategy for continual learning is to weigh a Bayesian ensemble of fixed functions. This suggests that if a (single) neural network could be interpreted as an ensemble, one could design effective algorithms that learn without forgetting. To realize this possibility, we observe that a neural network classifier with N parameters can be interpreted as a weighted ensemble of N classifiers, and that in the lazy regime limit these classifiers are fixed throughout learning. We call these classifiers the neural tangent experts and show they output valid probability distributions over the labels. We then derive the likelihood and posterior probability of each expert given past data. Surprisingly, the posterior updates for these experts are equivalent to a scaled and projected form of stochastic gradient descent (SGD) over the network weights. Away from the lazy regime, networks can be seen as ensembles of adaptive experts which improve over time. These results offer a new interpretation of neural networks as Bayesian ensembles of experts, providing a principled framework for understanding and mitigating catastrophic forgetting in continual learning settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext be89c744-0201-4351-85c2-0db3721f70eaCited by top-tier papers2
- On the Theory of Continual Learning with Gradient Descent for Neural NetworksHossein Taheri, Avishek Ghosh, Arya MazumdarICML 2026 · 2 citations
- Towards Understanding Catastrophic Forgetting in Two-layer Convolutional Neural NetworksBoqi Li, Youjun Wang, Weiwei LiuICML 2025
Builds on14
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du et al.NeurIPS 2022 · 933 citations
- How Good is the Bayes Posterior in Deep Neural Networks Really?Florian Wenzel, Kevin Roth, Bastiaan S. Veeling, Jakub Swiatkowski et al.ICML 2020 · 409 citations
- Understanding the Role of Training Regimes in Continual LearningSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Razvan Pascanu, Hassan GhasemzadehNeurIPS 2020 · 295 citations
- Effect of scale on catastrophic forgetting in neural networksVinay Venkatesh Ramasesh, Aitor Lewkowycz, Ethan DyerICLR 2022 · 212 citations
Related papers
- Natural continual learning: success is a journey, not (just) a destinationTa-Chu Kao, Kristopher T. Jensen, Gido van de Ven, Alberto Bernacchia et al.NeurIPS 2021 · 72 citations
- Learning curves for continual learning in neural networks: Self-knowledge transfer and forgettingRyo Karakida, Shotaro AkahoICLR 2022 · 16 citations
- Artificial Neuronal Ensembles with Learned Context Dependent GatingMatthew J. Tilley, Michelle Miller, David FreedmanICLR 2023 · 2 citations
- Kolmogorov-Arnold Networks Still Catastrophically Forget but Differently from MLPAnton Lee, Heitor Murilo Gomes, Yaqian Zhang, W. Bastiaan KleijnAAAI 2025 · 2 citations
- Learning to Continually Learn with the Bayesian PrincipleSoochan Lee, Hyeonseong Jeon, Jaehyeon Son, Gunhee KimICML 2024 · 11 citations
