On Monotonic Linear Interpolation of Neural Network Parameters
James Lucas, Juhan Bae, Michael R. Zhang, Stanislav Fort, Richard S. Zemel, Roger B. Grosse
摘要
Linearly interpolating between initial neural network parameters and converged parameters after training with SGD typically leads to a monotonic decrease in the training objective. This Monotonic Linear Interpolation (MLI) property, first observed by Goodfellow et al. [11], persists in spite of the non-convex objectives and highly non-linear training dynamics of neural networks. Extending on this work, we show that this property holds under varying network architectures, optimizers, and learning problems. We evaluate several possible hypotheses for this property that, to our knowledge, have not yet been explored. Additionally, we show that networks violating this property can be produced systematically, by forcing the weights to move far from initialization. The MLI property raises important questions about the loss landscape geometry of neural nets and highlights the need to further study its global properties.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Towards Scalable and Versatile Weight Space LearningKonstantin Schürholt, Michael W. Mahoney, Damian BorthICML 2024 · 被引用 39 次
- Git Re-Basin: Merging Models modulo Permutation SymmetriesSamuel K. Ainsworth, Jonathan Hayase, Siddhartha S. SrinivasaICLR 2023 · 被引用 32 次
- What Can Linear Interpolation of Neural Network Loss Landscapes Tell Us?Tiffany J. Vlaar, Jonathan FrankleICML 2022 · 被引用 31 次
- Measuring and Controlling Solution Degeneracy across Task-Trained Recurrent Neural NetworksAnn Huang, Satpreet Harcharan Singh, Flavio Martinelli, Kanaka RajanNeurIPS 2025 · 被引用 22 次
- Transferring Learning Trajectories of Neural NetworksDaiki ChijiwaICLR 2024 · 被引用 4 次
它引用的顶会 Paper4
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 被引用 750 次
- Proving the Lottery Ticket Hypothesis: Pruning is All You NeedEran Malach, Gilad Yehudai, Shai Shalev-Shwartz, Ohad ShamirICML 2020 · 被引用 327 次
- Landscape Connectivity and Dropout Stability of SGD Solutions for Over-parameterized Neural NetworksAlexander Shevchenko, Marco MondelliICML 2020 · 被引用 41 次
- When does preconditioning help or hurt generalization?Shun-ichi Amari, Jimmy Ba, Roger Baker Grosse, Xuechen Li 等ICLR 2021 · 被引用 11 次
相关 Paper
- Plateau in Monotonic Linear Interpolation - A "Biased" View of Loss Landscape for Deep NetworksXiang Wang, Annie N. Wang, Mo Zhou, Rong GeICLR 2023
- MLI Formula: A Nearly Scale-Invariant Solution with Noise PerturbationBowen Tao, Xin-Chun Li, De-Chuan ZhanICML 2024 · 被引用 1 次
- Chaotic Dynamics are Intrinsic to Neural Network Training with SGDLuis Herrmann, Maximilian Granz, Tim LandgrafNeurIPS 2022 · 被引用 15 次
- No Wrong Turns: The Simple Geometry Of Neural Networks Optimization PathsCharles Guille-Escuret, Hiroki Naganuma, Kilian Fatras, Ioannis MitliagkasICML 2024 · 被引用 9 次
- Flat Channels to Infinity in Neural Loss LandscapesFlavio Martinelli, Alexander van Meegen, Berfin Simsek, Wulfram Gerstner 等NeurIPS 2025 · 被引用 6 次
