Learning Continually by Spectral Regularization
Alex Lewandowski, Michal Bortkiewicz, Saurabh Kumar, András György, Dale Schuurmans, Mateusz Ostaszewski, Marlos C. Machado
摘要
Loss of plasticity is a phenomenon where neural networks can become more difficult to train over the course of learning. Continual learning algorithms seek to mitigate this effect by sustaining good performance while maintaining network trainability. We develop a new technique for improving continual learning inspired by the observation that the singular values of the neural network parameters at initialization are an important factor for trainability during early phases of learning. From this perspective, we derive a new spectral regularizer for continual learning that better sustains these beneficial initialization properties throughout training. In particular, the regularizer keeps the maximum singular value of each layer close to one. Spectral regularization directly ensures that gradient diversity is maintained throughout training, which promotes continual trainability, while minimally interfering with performance in a single task. We present an experimental analysis that shows how the proposed spectral regularizer can sustain trainability and performance across a range of model architectures in continual supervised and reinforcement learning settings. Spectral regularization is less sensitive to hyperparameters while demonstrating better training in individual tasks, sustaining trainability as new tasks arrive, and achieving better generalization performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Towards Higher Effective Rank in Parameter-Efficient Fine-Tuning Using Khatri-Rao ProductPaul Albert, Frederic Z. Zhang, Hemanth Saratchandran, Anton van den Hengel 等ICCV 2025 · 被引用 14 次
- FIRE: Frobenius-Isometry Reinitialization for Balancing the Stability-Plasticity TradeoffIsaac Han, Sangyeon Park, Seungwon Oh, Donghu Kim 等ICLR 2026 · 被引用 7 次
- Recurrent Self-Attention Dynamics: An Energy-Agnostic Perspective from JacobiansAkiyoshi Tomihari, Ryo KarakidaNeurIPS 2025 · 被引用 5 次
- SPHERE: Mitigating the Loss of Spectral Plasticity in Mixture-of-Experts for Deep Reinforcement LearningLirui Luo, Guoxi Zhang, Hongming Xu, Cong Fang 等ICML 2026 · 被引用 2 次
- Preserving Plasticity in Continual Learning via Dynamical IsometryAndries Rosseau, Robert Müller, Ann NoweICML 2026 · 被引用 1 次
它引用的顶会 Paper19
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya 等NeurIPS 2022 · 被引用 566 次
- Efficient Online Reinforcement Learning with Offline DataPhilip J. Ball, Laura Smith, Ilya Kostrikov, Sergey LevineICML 2023 · 被引用 326 次
- On Warm-Starting Neural Network TrainingJordan T. Ash, Ryan P. AdamsNeurIPS 2020 · 被引用 288 次
- The Primacy Bias in Deep Reinforcement LearningEvgenii Nikishin, Max Schwarzer, Pierluca D'Oro, Pierre-Luc Bacon 等ICML 2022 · 被引用 269 次
相关 Paper
- Spectral Collapse Drives Loss of Plasticity in Deep Continual LearningArjun Prakash, Naicheng He, Kaicheng Guo, Saket Tiwari 等ICML 2026
- Parseval Regularization for Continual Reinforcement LearningWesley Chung, Lynn Cherif, Doina Precup, David MegerNeurIPS 2024 · 被引用 27 次
- Self-Normalized Resets for Plasticity in Continual LearningVivek F. Farias, Adam Daniel JozefiakICLR 2025
- Plastic Learning with Deep Fourier FeaturesAlex Lewandowski, Dale Schuurmans, Marlos C. MachadoICLR 2025
- Understanding the Role of Training Regimes in Continual LearningSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Razvan Pascanu, Hassan GhasemzadehNeurIPS 2020 · 被引用 295 次
