Plateau in Monotonic Linear Interpolation - A "Biased" View of Loss Landscape for Deep Networks
Xiang Wang, Annie N. Wang, Mo Zhou, Rong Ge
Abstract
Monotonic linear interpolation (MLI) -on the line connecting a random initialization with the minimizer it converges to, the loss and accuracy are monotonic -is a phenomenon that is commonly observed in the training of neural networks. Such a phenomenon may seem to suggest that optimization of neural networks is easy. In this paper, we show that the MLI property is not necessarily related to the hardness of optimization problems, and empirical observations on MLI for deep neural networks depend heavily on the biases. In particular, we show that interpolating both weights and biases linearly leads to very different influences on the final output, and when different classes have different last-layer biases on a deep network, there will be a long plateau in both the loss and accuracy interpolation (which existing theory of MLI cannot explain). We also show how the last-layer biases for different classes can be different even on a perfectly balanced dataset using a simple model. Empirically we demonstrate that similar intuitions hold on practical networks and realistic datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b281a0b9-2189-46aa-aac9-8bb86d2e4b17Cited by top-tier papers3
- The Empirical Impact of Neural Parameter Symmetries, or Lack ThereofDerek Lim, Theo (Moe) Putterman, Robin Walters, Haggai Maron et al.NeurIPS 2024 · 25 citations
- Transferring Learning Trajectories of Neural NetworksDaiki ChijiwaICLR 2024 · 4 citations
- MLI Formula: A Nearly Scale-Invariant Solution with Noise PerturbationBowen Tao, Xin-Chun Li, De-Chuan ZhanICML 2024 · 1 citation
Builds on6
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent KernelStanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani et al.NeurIPS 2020 · 255 citations
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 151 citations
- Landscape Connectivity and Dropout Stability of SGD Solutions for Over-parameterized Neural NetworksAlexander Shevchenko, Marco MondelliICML 2020 · 41 citations
- Understanding Deflation Process in Over-parametrized Tensor DecompositionRong Ge, Yunwei Ren, Xiang Wang, Mo ZhouNeurIPS 2021 · 22 citations
Related papers
- On Monotonic Linear Interpolation of Neural Network ParametersJames Lucas, Juhan Bae, Michael R. Zhang, Stanislav Fort et al.ICML 2021 · 11 citations
- What Can Linear Interpolation of Neural Network Loss Landscapes Tell Us?Tiffany J. Vlaar, Jonathan FrankleICML 2022 · 31 citations
- Unique Properties of Flat Minima in Deep NetworksRotem Mulayoff, Tomer MichaeliICML 2020 · 43 citations
- Training invariances and the low-rank phenomenon: beyond linear networksThien Le, Stefanie JegelkaICLR 2022 · 39 citations
- On the Explicit Role of Initialization on the Convergence and Implicit Bias of Overparametrized Linear NetworksHancheng Min, Salma Tarmoun, René Vidal, Enrique MalladaICML 2021 · 53 citations
