Plateau in Monotonic Linear Interpolation - A "Biased" View of Loss Landscape for Deep Networks
Xiang Wang, Annie N. Wang, Mo Zhou, Rong Ge
摘要
Monotonic linear interpolation (MLI) -on the line connecting a random initialization with the minimizer it converges to, the loss and accuracy are monotonic -is a phenomenon that is commonly observed in the training of neural networks. Such a phenomenon may seem to suggest that optimization of neural networks is easy. In this paper, we show that the MLI property is not necessarily related to the hardness of optimization problems, and empirical observations on MLI for deep neural networks depend heavily on the biases. In particular, we show that interpolating both weights and biases linearly leads to very different influences on the final output, and when different classes have different last-layer biases on a deep network, there will be a long plateau in both the loss and accuracy interpolation (which existing theory of MLI cannot explain). We also show how the last-layer biases for different classes can be different even on a perfectly balanced dataset using a simple model. Empirically we demonstrate that similar intuitions hold on practical networks and realistic datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- The Empirical Impact of Neural Parameter Symmetries, or Lack ThereofDerek Lim, Theo (Moe) Putterman, Robin Walters, Haggai Maron 等NeurIPS 2024 · 被引用 25 次
- Transferring Learning Trajectories of Neural NetworksDaiki ChijiwaICLR 2024 · 被引用 4 次
- MLI Formula: A Nearly Scale-Invariant Solution with Noise PerturbationBowen Tao, Xin-Chun Li, De-Chuan ZhanICML 2024 · 被引用 1 次
它引用的顶会 Paper6
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 被引用 750 次
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent KernelStanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani 等NeurIPS 2020 · 被引用 255 次
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 被引用 151 次
- Landscape Connectivity and Dropout Stability of SGD Solutions for Over-parameterized Neural NetworksAlexander Shevchenko, Marco MondelliICML 2020 · 被引用 41 次
- Understanding Deflation Process in Over-parametrized Tensor DecompositionRong Ge, Yunwei Ren, Xiang Wang, Mo ZhouNeurIPS 2021 · 被引用 22 次
相关 Paper
- On Monotonic Linear Interpolation of Neural Network ParametersJames Lucas, Juhan Bae, Michael R. Zhang, Stanislav Fort 等ICML 2021 · 被引用 11 次
- What Can Linear Interpolation of Neural Network Loss Landscapes Tell Us?Tiffany J. Vlaar, Jonathan FrankleICML 2022 · 被引用 31 次
- Unique Properties of Flat Minima in Deep NetworksRotem Mulayoff, Tomer MichaeliICML 2020 · 被引用 43 次
- Training invariances and the low-rank phenomenon: beyond linear networksThien Le, Stefanie JegelkaICLR 2022 · 被引用 39 次
- On the Explicit Role of Initialization on the Convergence and Implicit Bias of Overparametrized Linear NetworksHancheng Min, Salma Tarmoun, René Vidal, Enrique MalladaICML 2021 · 被引用 53 次
