On the generalization of learning algorithms that do not converge
Nisha Chandramoorthy, Andreas Loukas, Khashayar Gatmiry, Stefanie Jegelka
摘要
Generalization analyses of deep learning typically assume that the training converges to a fixed point. But, recent results indicate that in practice, the weights of deep neural networks optimized with stochastic gradient descent often oscillate indefinitely. To reduce this discrepancy between theory and practice, this paper focuses on the generalization of neural networks whose training dynamics do not necessarily converge to fixed points. Our main contribution is to propose a notion of statistical algorithmic stability (SAS) that extends classical algorithmic stability to non-convergent algorithms and to study its connection to generalization. This ergodic-theoretic approach leads to new insights when compared to the traditional optimization and learning theory perspectives. We prove that the stability of the time-asymptotic behavior of a learning algorithm relates to its generalization and empirically demonstrate how loss dynamics can provide clues to generalization performance. Our findings provide evidence that networks that "train stably generalize better" even when the training continues indefinitely and the weights do not converge.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Generalization Bounds using Data-Dependent Fractal DimensionsBenjamin Dupuis, George Deligiannidis, Umut SimsekliICML 2023 · 被引用 17 次
- Identifying Equivalent Training DynamicsWilliam T. Redman, Juan M. Bello-Rivas, Maria Fonoberova, Ryan Mohr 等NeurIPS 2024 · 被引用 15 次
- Learning Trajectories are Generalization IndicatorsJingwen Fu, Zhizheng Zhang, Dacheng Yin, Yan Lu 等NeurIPS 2023 · 被引用 6 次
- Dynamic Momentum Recalibration in Online Gradient LearningZhipeng Yao, Rui Yu, Guisong Chang, Ying Li 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper8
- Understanding the unstable convergence of gradient descentKwangjun Ahn, Jingzhao Zhang, Suvrit SraICML 2022 · 被引用 89 次
- Optimizing Neural Networks via Koopman Operator TheoryAkshunna S. Dogra, William T. RedmanNeurIPS 2020 · 被引用 65 次
- Large Learning Rate Tames Homogeneity: Convergence and Balancing EffectYuqing Wang, Minshuo Chen, Tuo Zhao, Molei TaoICLR 2022 · 被引用 53 次
- Stochasticity of Deterministic Gradient Descent: Large Learning Rate for Multiscale Objective FunctionLingkai Kong, Molei TaoNeurIPS 2020 · 被引用 35 次
- On the Periodic Behavior of Neural Network Training with Batch Normalization and Weight DecayEkaterina Lobacheva, Maxim Kodryan, Nadezhda Chirkova, Andrey Malinin 等NeurIPS 2021 · 被引用 30 次
相关 Paper
- Neural Network Weights Do Not Converge to Stationary Points: An Invariant Measure PerspectiveJingzhao Zhang, Haochuan Li, Suvrit Sra, Ali JadbabaieICML 2022 · 被引用 14 次
- Uniform-in-Time Wasserstein Stability Bounds for (Noisy) Stochastic Gradient DescentLingjiong Zhu, Mert Gürbüzbalaban, Anant Raj, Umut SimsekliNeurIPS 2023 · 被引用 10 次
- Stability and Generalization Analysis of Gradient Methods for Shallow Neural NetworksYunwen Lei, Rong Jin, Yiming YingNeurIPS 2022 · 被引用 30 次
- Chaotic Dynamics are Intrinsic to Neural Network Training with SGDLuis Herrmann, Maximilian Granz, Tim LandgrafNeurIPS 2022 · 被引用 15 次
- Fractal Structure and Generalization Properties of Stochastic Optimization AlgorithmsAlexander Camuto, George Deligiannidis, Murat A. Erdogdu, Mert Gürbüzbalaban 等NeurIPS 2021 · 被引用 34 次
