Neural networks with late-phase weights
Johannes von Oswald, Seijin Kobayashi, João Sacramento, Alexander Meulemans, Christian Henning, Benjamin F. Grewe
Abstract
The largely successful method of training neural networks is to learn their weights using some variant of stochastic gradient descent (SGD). Here, we show that the solutions found by SGD can be further improved by ensembling a subset of the weights in late stages of learning. At the end of learning, we obtain back a single model by taking a spatial average in weight space. To avoid incurring increased computational costs, we investigate a family of low-dimensional late-phase weight models which interact multiplicatively with the remaining parameters. Our results show that augmenting standard models with late-phase weights improves generalization in established benchmarks such as CIFAR-10/100, ImageNet and enwik8. These findings are complemented with a theoretical analysis of a noisy quadratic problem which provides a simplified picture of the late phases of neural network learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2d1a8cfc-c8b2-4fce-abc6-1b1f728e42edCited by top-tier papers13
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- ZipIt! Merging Models from Different Tasks without TrainingGeorge Stoica, Daniel Bolya, Jakob Bjorner, Pratik Ramesh et al.ICLR 2024 · 185 citations
- Repulsive Deep Ensembles are BayesianFrancesco D'Angelo, Vincent FortuinNeurIPS 2021 · 141 citations
- Learning Neural Network SubspacesMitchell Wortsman, Maxwell Horton, Carlos Guestrin, Ali Farhadi et al.ICML 2021 · 101 citations
- Posterior Meta-Replay for Continual LearningChristian Henning, Maria R. Cervera, Francesco D'Angelo, Johannes von Oswald et al.NeurIPS 2021 · 78 citations
Builds on7
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong LearningYeming Wen, Dustin Tran, Jimmy BaICLR 2020 · 569 citations
- Continual learning with hypernetworksJohannes von Oswald, Christian Henning, João Sacramento, Benjamin F. GreweICLR 2020 · 412 citations
- Meta-Learning with Warped Gradient DescentSebastian Flennerhag, Andrei A. Rusu, Razvan Pascanu, Francesco Visin et al.ICLR 2020 · 221 citations
- Training BatchNorm and Only BatchNorm: On the Expressive Power of Random Features in CNNsJonathan Frankle, David J. Schwab, Ari S. MorcosICLR 2021 · 163 citations
Related papers
- Trainable Weight Averaging: Efficient Training by Optimizing Historical SolutionsTao Li, Zhehao Huang, Qinghua Tao, Yingwen Wu et al.ICLR 2023
- Stochastic Weight Averaging in Parallel: Large-Batch Training That Generalizes WellVipul Gupta, Santiago Akle Serrano, Dennis DeCosteICLR 2020 · 78 citations
- Lookaround Optimizer: k steps around, 1 step averageJiangtao Zhang, Shunyu Liu, Jie Song, Tongtian Zhu et al.NeurIPS 2023 · 12 citations
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 845 citations
- Stochastic Collapse: How Gradient Noise Attracts SGD Dynamics Towards Simpler SubnetworksFeng Chen, Daniel Kunin, Atsushi Yamamura, Surya GanguliNeurIPS 2023 · 52 citations
