Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methods
Taiji Suzuki, Shunta Akiyama
Abstract
Establishing a theoretical analysis that explains why deep learning can outperform shallow learning such as kernel methods is one of the biggest issues in the deep learning literature. Towards answering this question, we evaluate excess risk of a deep learning estimator trained by a noisy gradient descent with ridge regularization on a mildly overparameterized neural network, and discuss its superiority to a class of linear estimators that includes neural tangent kernel approach, random feature model, other kernel methods, -NN estimator and so on. We consider a teacher-student regression model, and eventually show that any linear estimator can be outperformed by deep learning in a sense of the minimax optimal rate especially for a high dimension setting. The obtained excess bounds are so-called fast learning rate which is faster than that is obtained by usual Rademacher complexity analysis. This discrepancy is induced by the non-convex geometry of the model and the noisy gradient descent used for neural network training provably reaches a near global optimal solution even though the loss landscape is highly non-convex. Although the noisy gradient descent does not employ any explicit or implicit sparsity inducing regularization, it shows a preferable generalization performance that dominates linear estimators.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cb8ef051-cbaf-447b-84f3-53af696a162dCited by top-tier papers10
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the RepresentationJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang et al.NeurIPS 2022 · 173 citations
- Classifying high-dimensional Gaussian mixtures: Where kernel methods fail and neural networks succeedMaria Refinetti, Sebastian Goldt, Florent Krzakala, Lenka ZdeborováICML 2021 · 83 citations
- Deep learning is adaptive to intrinsic dimensionality of model smoothness in anisotropic Besov spaceTaiji Suzuki, Atsushi NitandaNeurIPS 2021 · 76 citations
- Machine Learning For Elliptic PDEs: Fast Rate Generalization Bound, Neural Scaling Law and Minimax OptimalityYiping Lu, Haoxuan Chen, Jianfeng Lu, Lexing Ying et al.ICLR 2022 · 54 citations
- Stability & Generalisation of Gradient Descent for Shallow Neural Networks without the Neural Tangent KernelDominic Richards, Ilja KuzborskijNeurIPS 2021 · 43 citations
Builds on4
- Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural NetworksYu Bai, Jason D. LeeICLR 2020 · 128 citations
- Deep learning is adaptive to intrinsic dimensionality of model smoothness in anisotropic Besov spaceTaiji Suzuki, Atsushi NitandaNeurIPS 2021 · 76 citations
- Towards Understanding Hierarchical Learning: Benefits of Neural RepresentationsMinshuo Chen, Yu Bai, Jason D. Lee, Tuo Zhao et al.NeurIPS 2020 · 61 citations
- Generalization bound of globally optimal non-convex neural network training: Transportation map estimation by infinite dimensional Langevin dynamicsTaiji SuzukiNeurIPS 2020 · 25 citations
Related papers
- Excess Risk of Two-Layer ReLU Neural Networks in Teacher-Student Settings and its Superiority to Kernel MethodsShunta Akiyama, Taiji SuzukiICLR 2023 · 1 citation
- A Non-Parametric Regression Viewpoint : Generalization of Overparametrized Deep RELU Network Under Noisy ObservationsNamjoon Suh, Hyunouk Ko, Xiaoming HuoICLR 2022 · 15 citations
- The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of GeneralizationBen Adlam, Jeffrey PenningtonICML 2020 · 133 citations
- A Generalized Neural Tangent Kernel Analysis for Two-layer Neural NetworksZixiang Chen, Yuan Cao, Quanquan Gu, Tong ZhangNeurIPS 2020 · 82 citations
- On the Asymptotic Learning Curves of Kernel Ridge Regression under Power-law DecayYicheng Li, Haobo Zhang, Qian LinNeurIPS 2023 · 23 citations
