Towards Understanding Generalization via Decomposing Excess Risk Dynamics
Jiaye Teng, Jianhao Ma, Yang Yuan
Abstract
Generalization is one of the fundamental issues in machine learning. However, traditional techniques like uniform convergence may be unable to explain generalization under overparameterization (Nagarajan & Kolter, 2019) . As alternative approaches, techniques based on stability analyze the training dynamics and derive algorithm-dependent generalization bounds. Unfortunately, the stability-based bounds are still far from explaining the surprising generalization in deep learning since neural networks usually suffer from unsatisfactory stability. This paper proposes a novel decomposition framework to improve the stability-based bounds via a more fine-grained analysis of the signal and noise, inspired by the observation that neural networks converge relatively slowly when fitting noise (which indicates better stability). Concretely, we decompose the excess risk dynamics and apply the stability-based bound only on the noise component. The decomposition framework performs well in both linear regimes (overparameterized linear regression) and non-linear regimes (diagonal matrix recovery). Experiments on neural networks verify the utility of the decomposition framework.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0f7afbd2-7ea7-4b21-a76c-a0f0ed08eb55Cited by top-tier papers2
- Towards Data-Algorithm Dependent Generalization: a Case Study on Overparameterized Linear RegressionJing Xu, Jiaye Teng, Yang Yuan, Andrew C. YaoNeurIPS 2023 · 3 citations
- Finding Generalization Measures by Contrasting Signal and NoiseJiaye Teng, Bohang Zhang, Ruichen Li, Haowei He et al.ICML 2023
Builds on7
- Stability of Stochastic Gradient Descent on Nonsmooth Convex LossesRaef Bassily, Vitaly Feldman, Cristóbal Guzmán, Kunal TalwarNeurIPS 2020 · 240 citations
- Fine-Grained Analysis of Stability and Generalization for Stochastic Gradient DescentYunwen Lei, Yiming YingICML 2020 · 165 citations
- Sharpened Generalization Bounds based on Conditional Mutual Information and an Application to Noisy, Iterative AlgorithmsMahdi Haghifam, Jeffrey Negrea, Ashish Khisti, Daniel M. Roy et al.NeurIPS 2020 · 124 citations
- Understanding Double Descent Requires A Fine-Grained Bias-Variance DecompositionBen Adlam, Jeffrey PenningtonNeurIPS 2020 · 111 citations
- On Generalization Error Bounds of Noisy Gradient Methods for Non-Convex LearningJian Li, Xuanyuan Luo, Mingda QiaoICLR 2020 · 95 citations
Related papers
- Generalizability of Neural Networks Minimizing Empirical Risk Based on Expressive PowerLijia Yu, Yibo Miao, Yifan Zhu, Xiao-Shan Gao et al.ICLR 2025
- Stability and Generalization Analysis of Gradient Methods for Shallow Neural NetworksYunwen Lei, Rong Jin, Yiming YingNeurIPS 2022 · 30 citations
- Generalization bound of globally optimal non-convex neural network training: Transportation map estimation by infinite dimensional Langevin dynamicsTaiji SuzukiNeurIPS 2020 · 25 citations
- Compression based bound for non-compressed network: unified generalization error analysis of large compressible deep neural networkTaiji Suzuki, Hiroshi Abe, Tomoaki NishimuraICLR 2020 · 57 citations
- Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methodsTaiji Suzuki, Shunta AkiyamaICLR 2021 · 12 citations
