Approximate Heavy Tails in Offline (Multi-Pass) Stochastic Gradient Descent
Kruno Lehman, Alain Durmus, Umut Simsekli
Abstract
A recent line of empirical studies has demonstrated that SGD might exhibit a heavy-tailed behavior in practical settings, and the heaviness of the tails might correlate with the overall performance. In this paper, we investigate the emergence of such heavy tails. Previous works on this problem only considered, up to our knowledge, online (also called single-pass) SGD, in which the emergence of heavy tails in theoretical findings is contingent upon access to an infinite amount of data. Hence, the underlying mechanism generating the reported heavy-tailed behavior in practical settings, where the amount of training data is finite, is still not well-understood. Our contribution aims to fill this gap. In particular, we show that the stationary distribution of offline (also called multi-pass) SGD exhibits 'approximate' power-law tails and the approximation error is controlled by how fast the empirical distribution of the training data converges to the true underlying data distribution in the Wasserstein metric. Our main takeaway is that, as the number of data points increases, offline SGD will behave increasingly 'power-law-like'. To achieve this result, we first prove nonasymptotic Wasserstein convergence bounds for offline SGD to online SGD as the number of data points increases, which can be interesting on their own. Finally, we illustrate our theory on various experiments conducted on synthetic data and neural networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 51f236ec-e7bb-4b18-b647-3856c9b8e7fcCited by top-tier papers4
- What is the Long-Run Distribution of Stochastic Gradient Descent? A Large Deviations AnalysisWaïss Azizian, Franck Iutzeler, Jérôme Malick, Panayotis MertikopoulosICML 2024 · 17 citations
- Generalization Bounds for Heavy-Tailed SDEs through the Fractional Fokker-Planck EquationBenjamin Dupuis, Umut SimsekliICML 2024 · 6 citations
- The Global Convergence Time of Stochastic Gradient Descent in Non-Convex Landscapes: Sharp Estimates via Large DeviationsWaïss Azizian, Franck Iutzeler, Jérôme Malick, Panayotis MertikopoulosICML 2025
- Bregman meets Lévy: Stochastic Mirror Descent with Heavy-Tailed Noise in Continuous and Discrete TimePierre-Louis Cauvin, Panayotis MertikopoulosICML 2026
Builds on10
- Towards Theoretically Understanding Why Sgd Generalizes Better Than Adam in Deep LearningPan Zhou, Jiashi Feng, Chao Ma, Caiming Xiong et al.NeurIPS 2020 · 309 citations
- The Heavy-Tail Phenomenon in SGDMert Gürbüzbalaban, Umut Simsekli, Lingjiong ZhuICML 2021 · 165 citations
- Multiplicative Noise and Heavy Tails in Stochastic OptimizationLiam Hodgkinson, Michael W. MahoneyICML 2021 · 90 citations
- Hausdorff Dimension, Heavy Tails, and Generalization in Neural NetworksUmut Simsekli, Ozan Sener, George Deligiannidis, Murat A. ErdogduNeurIPS 2020 · 79 citations
- Heavy Tails in SGD and Compressibility of Overparametrized Neural NetworksMelih Barsbey, Milad Sefidgaran, Murat A. Erdogdu, Gaël Richard et al.NeurIPS 2021 · 57 citations
Related papers
- Algorithmic Stability of Heavy-Tailed SGD with General Loss FunctionsAnant Raj, Lingjiong Zhu, Mert Gürbüzbalaban, Umut SimsekliICML 2023 · 21 citations
- Emergence of heavy tails in homogenized stochastic gradient descentZhezhe Jiao, Martin Keller-ResselNeurIPS 2024 · 6 citations
- From Optimization to Generalization under Heavy-Tailed Data: The Role of Gradient ClippingAleksandr Shestakov, Martin Takac, Eduard GorbunovICML 2026
- High Probability Guarantees for Nonconvex Stochastic Gradient Descent with Heavy TailsShaojie Li, Yong LiuICML 2022 · 37 citations
- Error Analysis Affected by Heavy-Tailed Gradients for Non-Convex Pairwise Stochastic Gradient DescentJun Chen, Hong Chen, Bin Gu, Guodong Liu et al.AAAI 2025 · 1 citation
