Linear Mode Connectivity in Multitask and Continual Learning
Seyed-Iman Mirzadeh, Mehrdad Farajtabar, Dilan Görür, Razvan Pascanu, Hassan Ghasemzadeh
Abstract
Continual (sequential) training and multitask (simultaneous) training are often attempting to solve the same overall objective: to find a solution that performs well on all considered tasks. The main difference is in the training regimes, where continual learning can only have access to one task at a time, which for neural networks typically leads to catastrophic forgetting. That is, the solution found for a subsequent task does not perform well on the previous ones anymore. However, the relationship between the different minima that the two training regimes arrive at is not well understood. What sets them apart? Is there a local structure that could explain the difference in performance achieved by the two different schemes? Motivated by recent work showing that different minima of the same task are typically connected by very simple curves of low error, we investigate whether multitask and continual solutions are similarly connected. We empirically find that indeed such connectivity can be reliably achieved and, more interestingly, it can be done by a linear path, conditioned on having the same initialization for both. We thoroughly analyze this observation and discuss its significance for the continual learning process. Furthermore, we exploit this finding to propose an effective algorithm that constrains the sequentially learned minima to behave as the multitask solution. We show that our method outperforms several state of the art continual learning algorithms on various vision benchmarks 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 70dca3fb-7e8a-41b8-9ecd-0f622386eb6fCited by top-tier papers67
- Federated Continual Learning with Weighted Inter-client TransferJaehong Yoon, Wonyong Jeong, Giwoong Lee, Eunho Yang et al.ICML 2021 · 303 citations
- Patching open-vocabulary models by interpolating weightsGabriel Ilharco, Mitchell Wortsman, Samir Yitzhak Gadre, Shuran Song et al.NeurIPS 2022 · 230 citations
- Online Coreset Selection for Rehearsal-based Continual LearningJaehong Yoon, Divyam Madaan, Eunho Yang, Sung Ju HwangICLR 2022 · 181 citations
- Forget-free Continual Learning with Winning SubnetworksHaeyong Kang, Rusty John Lloyd Mina, Sultan Rizky Hikmawan Madjid, Jaehong Yoon et al.ICML 2022 · 159 citations
- Rehearsal revealed: The limits and merits of revisiting samples in continual learningEli Verwimp, Matthias De Lange, Tinne TuytelaarsICCV 2021 · 121 citations
Builds on7
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
- What is being transferred in transfer learning?Behnam Neyshabur, Hanie Sedghi, Chiyuan ZhangNeurIPS 2020 · 654 citations
- Supermasks in SuperpositionMitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi et al.NeurIPS 2020 · 364 citations
- Understanding the Role of Training Regimes in Continual LearningSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Razvan Pascanu, Hassan GhasemzadehNeurIPS 2020 · 295 citations
- Using Hindsight to Anchor Past Knowledge in Continual LearningArslan Chaudhry, Albert Gordo, Puneet K. Dokania, Philip H. S. Torr et al.AAAI 2021 · 279 citations
Related papers
- Learning curves for continual learning in neural networks: Self-knowledge transfer and forgettingRyo Karakida, Shotaro AkahoICLR 2022 · 16 citations
- Memory-Statistics Tradeoff in Continual Learning with Structural RegularizationHaoran Li, Jingfeng Wu, Vladimir BravermanICLR 2026 · 4 citations
- Learning without Prejudices: Continual Unbiased Learning via Benign and Malignant ForgettingMyeongho Jeon, Hyoje Lee, Yedarm Seong, Myungjoo KangICLR 2023
- Continual Learning in the Teacher-Student Setup: Impact of Task SimilaritySebastian Lee, Sebastian Goldt, Andrew M. SaxeICML 2021 · 98 citations
- CLR: Channel-wise Lightweight Reprogramming for Continual LearningYunhao Ge, Yuecheng Li, Shuo Ni, Jiaping Zhao et al.ICCV 2023 · 16 citations
