Hausdorff Dimension, Heavy Tails, and Generalization in Neural Networks
Umut Simsekli, Ozan Sener, George Deligiannidis, Murat A. Erdogdu
摘要
Despite its success in a wide range of applications, characterizing the generalization properties of stochastic gradient descent (SGD) in non-convex deep learning problems is still an important challenge. While modeling the trajectories of SGD via stochastic differential equations (SDE) under heavy-tailed gradient noise has recently shed light over several peculiar characteristics of SGD, a rigorous treatment of the generalization properties of such SDEs in a learning theoretical framework is still missing. Aiming to bridge this gap, in this paper, we prove generalization bounds for SGD under the assumption that its trajectories can be well-approximated by a Feller process, which defines a rich class of Markov processes that include several recent SDE representations (both Brownian or heavy-tailed) as its special case. We show that the generalization error can be controlled by the Hausdorff dimension of the trajectories, which is intimately linked to the tail behavior of the driving process. Our results imply that heavier-tailed processes should achieve better generalization; hence, the tail-index of the process can be used as a notion of "capacity metric". We support our theory with experiments on deep neural networks illustrating that the proposed capacity metric accurately estimates the generalization error, and it does not necessarily grow with the number of parameters unlike the existing capacity metrics in the literature. 34th Conference on Neural Information Processing Systems (NeurIPS 2020),
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper36
- Intrinsic Dimension, Persistent Homology and Generalization in Neural NetworksTolga Birdal, Aaron Lou, Leonidas J. Guibas, Umut SimsekliNeurIPS 2021 · 被引用 94 次
- Heavy Tails in SGD and Compressibility of Overparametrized Neural NetworksMelih Barsbey, Milad Sefidgaran, Murat A. Erdogdu, Gaël Richard 等NeurIPS 2021 · 被引用 57 次
- Noise and Fluctuation of Finite Learning Rate Stochastic Gradient DescentKangqiao Liu, Liu Ziyin, Masahito UedaICML 2021 · 被引用 46 次
- Fractal Structure and Generalization Properties of Stochastic Optimization AlgorithmsAlexander Camuto, George Deligiannidis, Murat A. Erdogdu, Mert Gürbüzbalaban 等NeurIPS 2021 · 被引用 34 次
- Temperature Balancing, Layer-wise Weight Analysis, and Neural Network TrainingYefan Zhou, Tianyu Pang, Keqin Liu, Charles H. Martin 等NeurIPS 2023 · 被引用 29 次
它引用的顶会 Paper3
- The Heavy-Tail Phenomenon in SGDMert Gürbüzbalaban, Umut Simsekli, Lingjiong ZhuICML 2021 · 被引用 165 次
- Multiplicative Noise and Heavy Tails in Stochastic OptimizationLiam Hodgkinson, Michael W. MahoneyICML 2021 · 被引用 90 次
- Fractional Underdamped Langevin Dynamics: Retargeting SGD with Momentum under Heavy-Tailed Gradient NoiseUmut Simsekli, Lingjiong Zhu, Yee Whye Teh, Mert GürbüzbalabanICML 2020 · 被引用 58 次
相关 Paper
- Algorithmic Stability of Heavy-Tailed SGD with General Loss FunctionsAnant Raj, Lingjiong Zhu, Mert Gürbüzbalaban, Umut SimsekliICML 2023 · 被引用 21 次
- Emergence of heavy tails in homogenized stochastic gradient descentZhezhe Jiao, Martin Keller-ResselNeurIPS 2024 · 被引用 6 次
- Learning Trajectories are Generalization IndicatorsJingwen Fu, Zhizheng Zhang, Dacheng Yin, Yan Lu 等NeurIPS 2023 · 被引用 6 次
- Robustness Analysis of Non-Convex Stochastic Gradient Descent using Biased ExpectationsKevin Scaman, Cédric MalherbeNeurIPS 2020 · 被引用 37 次
- Generalization Bounds for Heavy-Tailed SDEs through the Fractional Fokker-Planck EquationBenjamin Dupuis, Umut SimsekliICML 2024 · 被引用 6 次
