Stochasticity of Deterministic Gradient Descent: Large Learning Rate for Multiscale Objective Function
Lingkai Kong, Molei Tao
Abstract
This article suggests that deterministic Gradient Descent, which does not use any stochastic gradient approximation, can still exhibit stochastic behaviors. In particular, it shows that if the objective function exhibit multiscale behaviors, then in a large learning rate regime which only resolves the macroscopic but not the microscopic details of the objective, the deterministic GD dynamics can become chaotic and convergent not to a local minimizer but to a statistical distribution. In this sense, deterministic GD resembles stochastic GD even though no stochasticity is injected. A sufficient condition is also established for approximating this long-time statistical limit by a rescaled Gibbs distribution, which for example allows escapes from local minima to be quantified. Both theoretical and numerical demonstrations are provided, and the theoretical part relies on the construction of a stochastic map that uses bounded noise (as opposed to Gaussian noise).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc08d7f5-8d99-46b9-adf6-966d8a5f4b90Cited by top-tier papers13
- Understanding the Generalization Benefit of Normalization Layers: Sharpness ReductionKaifeng Lyu, Zhiyuan Li, Sanjeev AroraNeurIPS 2022 · 111 citations
- Stochastic Training is Not Necessary for GeneralizationJonas Geiping, Micah Goldblum, Phillip Pope, Michael Moeller et al.ICLR 2022 · 83 citations
- Large Learning Rate Tames Homogeneity: Convergence and Balancing EffectYuqing Wang, Minshuo Chen, Tuo Zhao, Molei TaoICLR 2022 · 53 citations
- Implicit Bias of Gradient Descent for Logistic Regression at the Edge of StabilityJingfeng Wu, Vladimir Braverman, Jason D. LeeNeurIPS 2023 · 46 citations
- Generalization Bounds using Lower Tail Exponents in Stochastic OptimizersLiam Hodgkinson, Umut Simsekli, Rajiv Khanna, Michael W. MahoneyICML 2022 · 29 citations
Builds on2
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell et al.NeurIPS 2020 · 4,008 citations
- Generalization Error Bounds of Gradient Descent for Learning Over-Parameterized Deep ReLU NetworksYuan Cao, Quanquan GuAAAI 2020 · 168 citations
Related papers
- What is the Long-Run Distribution of Stochastic Gradient Descent? A Large Deviations AnalysisWaïss Azizian, Franck Iutzeler, Jérôme Malick, Panayotis MertikopoulosICML 2024 · 17 citations
- Phase diagram of Stochastic Gradient Descent in high-dimensional two-layer neural networksRodrigo Veiga, Ludovic Stephan, Bruno Loureiro, Florent Krzakala et al.NeurIPS 2022 · 59 citations
- Chaotic Regularization and Heavy-Tailed Limits for Deterministic Gradient DescentSoon Hoe Lim, Yijun Wan, Umut SimsekliNeurIPS 2022 · 14 citations
- Dynamical mean-field theory for stochastic gradient descent in Gaussian mixture classificationFrancesca Mignacco, Florent Krzakala, Pierfrancesco Urbani, Lenka ZdeborováNeurIPS 2020 · 95 citations
- The Global Convergence Time of Stochastic Gradient Descent in Non-Convex Landscapes: Sharp Estimates via Large DeviationsWaïss Azizian, Franck Iutzeler, Jérôme Malick, Panayotis MertikopoulosICML 2025
