On Monotonic Linear Interpolation of Neural Network Parameters
James Lucas, Juhan Bae, Michael R. Zhang, Stanislav Fort, Richard S. Zemel, Roger B. Grosse
Abstract
Linearly interpolating between initial neural network parameters and converged parameters after training with SGD typically leads to a monotonic decrease in the training objective. This Monotonic Linear Interpolation (MLI) property, first observed by Goodfellow et al. [11], persists in spite of the non-convex objectives and highly non-linear training dynamics of neural networks. Extending on this work, we show that this property holds under varying network architectures, optimizers, and learning problems. We evaluate several possible hypotheses for this property that, to our knowledge, have not yet been explored. Additionally, we show that networks violating this property can be produced systematically, by forcing the weights to move far from initialization. The MLI property raises important questions about the loss landscape geometry of neural nets and highlights the need to further study its global properties.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext df73f603-345f-467f-b790-2322f374a4dfCited by top-tier papers5
- Towards Scalable and Versatile Weight Space LearningKonstantin Schürholt, Michael W. Mahoney, Damian BorthICML 2024 · 39 citations
- Git Re-Basin: Merging Models modulo Permutation SymmetriesSamuel K. Ainsworth, Jonathan Hayase, Siddhartha S. SrinivasaICLR 2023 · 32 citations
- What Can Linear Interpolation of Neural Network Loss Landscapes Tell Us?Tiffany J. Vlaar, Jonathan FrankleICML 2022 · 31 citations
- Measuring and Controlling Solution Degeneracy across Task-Trained Recurrent Neural NetworksAnn Huang, Satpreet Harcharan Singh, Flavio Martinelli, Kanaka RajanNeurIPS 2025 · 22 citations
- Transferring Learning Trajectories of Neural NetworksDaiki ChijiwaICLR 2024 · 4 citations
Builds on4
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
- Proving the Lottery Ticket Hypothesis: Pruning is All You NeedEran Malach, Gilad Yehudai, Shai Shalev-Shwartz, Ohad ShamirICML 2020 · 327 citations
- Landscape Connectivity and Dropout Stability of SGD Solutions for Over-parameterized Neural NetworksAlexander Shevchenko, Marco MondelliICML 2020 · 41 citations
- When does preconditioning help or hurt generalization?Shun-ichi Amari, Jimmy Ba, Roger Baker Grosse, Xuechen Li et al.ICLR 2021 · 11 citations
Related papers
- Plateau in Monotonic Linear Interpolation - A "Biased" View of Loss Landscape for Deep NetworksXiang Wang, Annie N. Wang, Mo Zhou, Rong GeICLR 2023
- MLI Formula: A Nearly Scale-Invariant Solution with Noise PerturbationBowen Tao, Xin-Chun Li, De-Chuan ZhanICML 2024 · 1 citation
- Chaotic Dynamics are Intrinsic to Neural Network Training with SGDLuis Herrmann, Maximilian Granz, Tim LandgrafNeurIPS 2022 · 15 citations
- No Wrong Turns: The Simple Geometry Of Neural Networks Optimization PathsCharles Guille-Escuret, Hiroki Naganuma, Kilian Fatras, Ioannis MitliagkasICML 2024 · 9 citations
- Flat Channels to Infinity in Neural Loss LandscapesFlavio Martinelli, Alexander van Meegen, Berfin Simsek, Wulfram Gerstner et al.NeurIPS 2025 · 6 citations
