Unique Properties of Flat Minima in Deep Networks
Rotem Mulayoff, Tomer Michaeli
Abstract
It is well known that (stochastic) gradient descent has an implicit bias towards flat minima. In deep neural network training, this mechanism serves to screen out minima. However, the precise effect that this has on the trained network is not yet fully understood. In this paper, we characterize the flat minima in linear neural networks trained with a quadratic loss. First, we show that linear ResNets with zero initialization necessarily converge to the flattest of all minima. We then prove that these minima correspond to nearly balanced networks whereby the gain from the input to any intermediate representation does not change drastically from one layer to the next. Finally, we show that consecutive layers in flat minima solutions are coupled. That is, one of the left singular vectors of each weight matrix, equals one of the right singular vectors of the next matrix. This forms a distinct path from input to output, that, as we show, is dedicated to the signal that experiences the largest gain end-to-end. Experiments indicate that these properties are characteristic of both linear and nonlinear models trained in practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f568d793-58c0-4e9b-a05a-7ba6f3a295dfCited by top-tier papers16
- Towards Understanding Sharpness-Aware MinimizationMaksym Andriushchenko, Nicolas FlammarionICML 2022 · 190 citations
- Understanding the Generalization Benefit of Normalization Layers: Sharpness ReductionKaifeng Lyu, Zhiyuan Li, Sanjeev AroraNeurIPS 2022 · 111 citations
- DR3: Value-Based Deep Reinforcement Learning Requires Explicit RegularizationAviral Kumar, Rishabh Agarwal, Tengyu Ma, Aaron C. Courville et al.ICLR 2022 · 85 citations
- The Implicit Bias of Minima Stability: A View from Function SpaceRotem Mulayoff, Tomer Michaeli, Daniel SoudryNeurIPS 2021 · 65 citations
- Gradient Descent on Neural Networks Typically Occurs at the Edge of StabilityJeremy Cohen, Simran Kaur, Yuanzhi Li, J. Zico Kolter et al.ICLR 2021 · 22 citations
Builds on1
Related papers
- On the spectral bias of two-layer linear networksAditya Vardhan Varre, Maria-Luiza Vladarean, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2023 · 27 citations
- On the Global Convergence of Training Deep Linear ResNetsDifan Zou, Philip M. Long, Quanquan GuICLR 2020 · 44 citations
- Deep linear networks for regression are implicitly regularized towards flat minimaPierre Marion, Lénaïc ChizatNeurIPS 2024 · 21 citations
- On the Explicit Role of Initialization on the Convergence and Implicit Bias of Overparametrized Linear NetworksHancheng Min, Salma Tarmoun, René Vidal, Enrique MalladaICML 2021 · 53 citations
- How much does Initialization Affect Generalization?Sameera Ramasinghe, Lachlan Ewen MacDonald, Moshiur R. Farazi, Hemanth Saratchandran et al.ICML 2023 · 9 citations
