Disentangling and mitigating the impact of task similarity for continual learning
Naoki Hiratani
摘要
Continual learning of partially similar tasks poses a challenge for artificial neural networks, as task similarity presents both an opportunity for knowledge transfer and a risk of interference and catastrophic forgetting. However, it remains unclear how task similarity in input features and readout patterns influences knowledge transfer and forgetting, as well as how they interact with common algorithms for continual learning. Here, we develop a linear teacher-student model with latent structure and show analytically that high input feature similarity coupled with low readout similarity is catastrophic for both knowledge transfer and retention. Conversely, the opposite scenario is relatively benign. Our analysis further reveals that task-dependent activity gating improves knowledge retention at the expense of transfer, while task-dependent plasticity gating does not affect either retention or transfer performance at the over-parameterized limit. In contrast, weight regularization based on the Fisher information metric significantly improves retention, regardless of task similarity, without compromising transfer performance. Nevertheless, its diagonal approximation and regularization in the Euclidean space are much less robust against task similarity. We demonstrate consistent results in a permuted MNIST task with latent variables. Overall, this work provides insights into when continual learning is difficult and how to mitigate it.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Fast Last-Iterate Convergence of SGD in the Smooth Interpolation RegimeAmit Attia, Matan Schliserman, Uri Sherman, Tomer KorenNeurIPS 2025 · 被引用 18 次
- Are Greedy Task Orderings Better Than Random in Continual Linear Regression?Matan Tsipory, Ran Levinstein, Itay Evron, Mark Kong 等NeurIPS 2025 · 被引用 5 次
- Autoencoder-Based Hybrid Replay for Class-Incremental LearningMilad Khademi Nori, Il-Min Kim, Guanghui WangICML 2025
- Optimal Task Order for Continual Learning of Multiple TasksZiyan Li, Naoki HirataniICML 2025
- Enhancing Continual Learning of Vision-Language Models via Dynamic Prefix WeightingHyeonseo Jang, Hyuk Kwon, Kibok LeeCVPR 2026
它引用的顶会 Paper13
- Continual learning in recurrent neural networksBenjamin Ehret, Christian Henning, Maria R. Cervera, Alexander Meulemans 等ICLR 2021 · 被引用 4,433 次
- Anatomy of Catastrophic Forgetting: Hidden Representations and Task SemanticsVinay Venkatesh Ramasesh, Ethan Dyer, Maithra RaghuICLR 2021 · 被引用 207 次
- Continual Learning of a Mixed Sequence of Similar and Dissimilar TasksZixuan Ke, Bing Liu, Xingchang HuangNeurIPS 2020 · 被引用 173 次
- Optimal Continual Learning has Perfect Memory and is NP-hardJeremias Knoblauch, Hisham Husain, Tom DietheICML 2020 · 被引用 116 次
- Continual Learning in the Teacher-Student Setup: Impact of Task SimilaritySebastian Lee, Sebastian Goldt, Andrew M. SaxeICML 2021 · 被引用 98 次
相关 Paper
- Learning curves for continual learning in neural networks: Self-knowledge transfer and forgettingRyo Karakida, Shotaro AkahoICLR 2022 · 被引用 16 次
- Artificial Neuronal Ensembles with Learned Context Dependent GatingMatthew J. Tilley, Michelle Miller, David FreedmanICLR 2023 · 被引用 2 次
- Conditional Channel Gated Networks for Task-Aware Continual LearningDavide Abati, Jakub M. Tomczak, Tijmen Blankevoort, Simone Calderara 等CVPR 2020
- Is Forgetting Less a Good Inductive Bias for Forward Transfer?Jiefeng Chen, Timothy Nguyen, Dilan Görür, Arslan ChaudhryICLR 2023 · 被引用 1 次
- Continual learning with hypernetworksJohannes von Oswald, Christian Henning, João Sacramento, Benjamin F. GreweICLR 2020 · 被引用 412 次
