Eliminating Sharp Minima from SGD with Truncated Heavy-tailed Noise
Xingyu Wang, Sewoong Oh, Chang-Han Rhee
摘要
The empirical success of deep learning is often attributed to SGD's mysterious ability to avoid sharp local minima in the loss landscape, as sharp minima are known to lead to poor generalization. Recently, empirical evidence of heavy-tailed gradient noise was reported in many deep learning tasks, and it was shown in Simsekli (2019a,b) that SGD can escape sharp local minima under the presence of such heavy-tailed gradient noise, providing a partial solution to the mystery. In this work, we analyze a popular variant of SGD where gradients are truncated above a fixed threshold. We show that it achieves a stronger notion of avoiding sharp minima: it can effectively eliminate sharp local minima entirely from its training trajectory. We characterize the dynamics of truncated SGD driven by heavy-tailed noises. First, we show that the truncation threshold and width of the attraction field dictate the order of the first exit time from the associated local minimum. Moreover, when the objective function satisfies appropriate structural conditions, we prove that as the learning rate decreases, the dynamics of heavy-tailed truncated SGD closely resemble those of a continuous-time Markov chain that never visits any sharp minima. Real data experiments on deep learning confirm our theoretical prediction that heavy-tailed SGD with gradient clipping finds a"flatter"local minima and achieves better generalization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- The alignment property of SGD noise and how it helps select flat minima: A stability analysisLei Wu, Mingze Wang, Weijie SuNeurIPS 2022 · 被引用 80 次
- Anticorrelated Noise Injection for Improved GeneralizationAntonio Orvieto, Hans Kersting, Frank Proske, Francis R. Bach 等ICML 2022 · 被引用 58 次
- Special Properties of Gradient Descent with Large Learning RatesAmirkeivan Mohtashami, Martin Jaggi, Sebastian U. StichICML 2023 · 被引用 16 次
- Zeroth-Order Optimization Finds Flat MinimaLiang Zhang, Bingcong Li, Kiran Koshy Thekumparampil, Sewoong Oh 等NeurIPS 2025 · 被引用 8 次
- Towards Understanding The Calibration Benefits of Sharpness-Aware MinimizationChengli Tan, Yubo Zhou, Haishan Ye, Guang Dai 等ICLR 2026 · 被引用 3 次
它引用的顶会 Paper6
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Implementation Matters in Deep RL: A Case Study on PPO and TRPOLogan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras 等ICLR 2020 · 被引用 305 次
- Stochastic Optimization with Heavy-Tailed Noise via Accelerated Gradient ClippingEduard Gorbunov, Marina Danilova, Alexander V. GasnikovNeurIPS 2020 · 被引用 181 次
- The Heavy-Tail Phenomenon in SGDMert Gürbüzbalaban, Umut Simsekli, Lingjiong ZhuICML 2021 · 被引用 165 次
- On Proximal Policy Optimization's Heavy-tailed GradientsSaurabh Garg, Joshua Zhanson, Emilio Parisotto, Adarsh Prasad 等ICML 2021 · 被引用 32 次
相关 Paper
- Fractional Underdamped Langevin Dynamics: Retargeting SGD with Momentum under Heavy-Tailed Gradient NoiseUmut Simsekli, Lingjiong Zhu, Yee Whye Teh, Mert GürbüzbalabanICML 2020 · 被引用 58 次
- Stability and Generalization of Nonconvex Optimization with Heavy-Tailed NoiseHongxu Chen, Ke Wei, Xiaoming Yuan, Luo LuoICML 2026
- Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-TailedSavelii Chezhegov, Yaroslav Klyukin, Andrei Semenov, Aleksandr Beznosikov 等ICML 2025
- Why are Adaptive Methods Good for Attention Models?Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim 等NeurIPS 2020 · 被引用 397 次
- Towards Theoretically Understanding Why Sgd Generalizes Better Than Adam in Deep LearningPan Zhou, Jiashi Feng, Chao Ma, Caiming Xiong 等NeurIPS 2020 · 被引用 309 次
