Fast Mixing of Stochastic Gradient Descent with Normalization and Weight Decay
Zhiyuan Li, Tianhao Wang, Dingli Yu
摘要
We prove the Fast Equilibrium Conjecture proposed by Li et al. [1], i.e. , stochastic gradient descent (SGD) on a scale-invariant loss ( e.g. , using networks with various normalization schemes) with learning rate ⌘ and weight decay factor � mixes in function space in e O (1 / ( ⌘� )) steps, under two standard assumptions: (1) the noise covariance matrix is non-degenerate and (2) the minimizers of the loss form a connected, compact and analytic manifold. The analysis uses the framework of Li et al. [2] and shows that for every T > 0 , the iterates of SGD with learning rate ⌘ and weight decay factor � on the scale-invariant loss converge in distribution in ln(1 + T � / ⌘ ) / (4 ⌘� ) iterations as ⌘� ! 0 while satisfying ⌘ O ( � ) O (1) . Moreover, the evolution of the limiting distribution can be described by a stochastic differential equation that mixes to the same equilibrium distribution for every initialization around the manifold of minimizers as T ! 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Why Do We Need Weight Decay in Modern Deep Learning?Francesco D'Angelo, Maksym Andriushchenko, Aditya Vardhan Varre, Nicolas FlammarionNeurIPS 2024 · 被引用 101 次
- Sharpness Minimization Algorithms Do Not Only Minimize Sharpness To Achieve Better GeneralizationKaiyue Wen, Zhiyuan Li, Tengyu MaNeurIPS 2023 · 被引用 53 次
- Implicit Bias of AdamW: ℓ∞-Norm Constrained OptimizationShuo Xie, Zhiyuan LiICML 2024 · 被引用 46 次
- Rotational Equilibrium: How Weight Decay Balances Learning Across Neural NetworksAtli Kosson, Bettina Messmer, Martin JaggiICML 2024 · 被引用 39 次
- What is the Long-Run Distribution of Stochastic Gradient Descent? A Large Deviations AnalysisWaïss Azizian, Franck Iutzeler, Jérôme Malick, Panayotis MertikopoulosICML 2024 · 被引用 17 次
它引用的顶会 Paper23
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 被引用 267 次
- Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep NetworksSoham De, Samuel L. SmithNeurIPS 2020 · 被引用 173 次
- Label Noise SGD Provably Prefers Flat Global MinimizersAlex Damian, Tengyu Ma, Jason D. LeeNeurIPS 2021 · 被引用 155 次
- Understanding Gradient Descent on the Edge of Stability in Deep LearningSanjeev Arora, Zhiyuan Li, Abhishek PanigrahiICML 2022 · 被引用 139 次
相关 Paper
- Fast Equilibrium of SGD in Generic SituationsZhiyuan Liu, Yi Wang, Zhiren WangICLR 2024 · 被引用 1 次
- Three-stage Evolution and Fast Equilibrium for SGD with Non-degerate Critical PointsYi Wang, Zhiren WangICML 2022 · 被引用 4 次
- Reconciling Modern Deep Learning with Traditional Optimization Analyses: The Intrinsic Learning RateZhiyuan Li, Kaifeng Lyu, Sanjeev AroraNeurIPS 2020 · 被引用 93 次
- On the Almost Sure Convergence of Stochastic Gradient Descent in Non-Convex ProblemsPanayotis Mertikopoulos, Nadav Hallak, Ali Kavis, Volkan CevherNeurIPS 2020 · 被引用 120 次
- Last iterate convergence of SGD for Least-Squares in the Interpolation regimeAditya Vardhan Varre, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2021 · 被引用 52 次
