Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methods
Taiji Suzuki, Shunta Akiyama
摘要
Establishing a theoretical analysis that explains why deep learning can outperform shallow learning such as kernel methods is one of the biggest issues in the deep learning literature. Towards answering this question, we evaluate excess risk of a deep learning estimator trained by a noisy gradient descent with ridge regularization on a mildly overparameterized neural network, and discuss its superiority to a class of linear estimators that includes neural tangent kernel approach, random feature model, other kernel methods, -NN estimator and so on. We consider a teacher-student regression model, and eventually show that any linear estimator can be outperformed by deep learning in a sense of the minimax optimal rate especially for a high dimension setting. The obtained excess bounds are so-called fast learning rate which is faster than that is obtained by usual Rademacher complexity analysis. This discrepancy is induced by the non-convex geometry of the model and the noisy gradient descent used for neural network training provably reaches a near global optimal solution even though the loss landscape is highly non-convex. Although the noisy gradient descent does not employ any explicit or implicit sparsity inducing regularization, it shows a preferable generalization performance that dominates linear estimators.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the RepresentationJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang 等NeurIPS 2022 · 被引用 173 次
- Classifying high-dimensional Gaussian mixtures: Where kernel methods fail and neural networks succeedMaria Refinetti, Sebastian Goldt, Florent Krzakala, Lenka ZdeborováICML 2021 · 被引用 83 次
- Deep learning is adaptive to intrinsic dimensionality of model smoothness in anisotropic Besov spaceTaiji Suzuki, Atsushi NitandaNeurIPS 2021 · 被引用 76 次
- Machine Learning For Elliptic PDEs: Fast Rate Generalization Bound, Neural Scaling Law and Minimax OptimalityYiping Lu, Haoxuan Chen, Jianfeng Lu, Lexing Ying 等ICLR 2022 · 被引用 54 次
- Stability & Generalisation of Gradient Descent for Shallow Neural Networks without the Neural Tangent KernelDominic Richards, Ilja KuzborskijNeurIPS 2021 · 被引用 43 次
它引用的顶会 Paper4
- Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural NetworksYu Bai, Jason D. LeeICLR 2020 · 被引用 128 次
- Deep learning is adaptive to intrinsic dimensionality of model smoothness in anisotropic Besov spaceTaiji Suzuki, Atsushi NitandaNeurIPS 2021 · 被引用 76 次
- Towards Understanding Hierarchical Learning: Benefits of Neural RepresentationsMinshuo Chen, Yu Bai, Jason D. Lee, Tuo Zhao 等NeurIPS 2020 · 被引用 61 次
- Generalization bound of globally optimal non-convex neural network training: Transportation map estimation by infinite dimensional Langevin dynamicsTaiji SuzukiNeurIPS 2020 · 被引用 25 次
相关 Paper
- Excess Risk of Two-Layer ReLU Neural Networks in Teacher-Student Settings and its Superiority to Kernel MethodsShunta Akiyama, Taiji SuzukiICLR 2023 · 被引用 1 次
- A Non-Parametric Regression Viewpoint : Generalization of Overparametrized Deep RELU Network Under Noisy ObservationsNamjoon Suh, Hyunouk Ko, Xiaoming HuoICLR 2022 · 被引用 15 次
- The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of GeneralizationBen Adlam, Jeffrey PenningtonICML 2020 · 被引用 133 次
- A Generalized Neural Tangent Kernel Analysis for Two-layer Neural NetworksZixiang Chen, Yuan Cao, Quanquan Gu, Tong ZhangNeurIPS 2020 · 被引用 82 次
- On the Asymptotic Learning Curves of Kernel Ridge Regression under Power-law DecayYicheng Li, Haobo Zhang, Qian LinNeurIPS 2023 · 被引用 23 次
