Generalization of Two-layer Neural Networks: An Asymptotic Viewpoint
Jimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Denny Wu, Tianzong Zhang
摘要
This paper investigates the generalization properties of two-layer neural networks in high-dimensions, i.e. when the number of samples n, features d, and neurons h tend to infinity at the same rate. Specifically, we derive the exact population risk of the unregularized least squares regression problem with two-layer neural networks when either the first or the second layer is trained using a gradient flow under different initialization setups. When only the second layer coefficients are optimized, we recover the double descent phenomenon: a cusp in the population risk appears at h ≈ n and further overparameterization decreases the risk. In contrast, when the first layer weights are optimized, we highlight how different scales of initialization lead to different inductive bias, and show that the resulting risk is independent of overparameterization. Our theoretical and experimental results suggest that previously studied model setups that provably give rise to double descent might not translate to optimizing two-layer neural networks. INTRODUCTION In modern neural networks, the number of parameters can easily exceed the number of training samples, yet in many circumstances, there is little sign of overfitting even in the absence of explicit regularization (Zhang et al., 2016) . This phenomenon is usually explained by the interplay between the model architecture and the optimization method. Existing works have analyzed the implicit regularization of gradient descent on simple models (
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- How Neural Networks Extrapolate: From Feedforward to Graph Neural NetworksKeyulu Xu, Mozhi Zhang, Jingling Li, Simon Shaolei Du 等ICLR 2021 · 被引用 364 次
- Rethinking Bias-Variance Trade-off for Generalization of Neural NetworksZitong Yang, Yaodong Yu, Chong You, Jacob Steinhardt 等ICML 2020 · 被引用 219 次
- How Does Knowledge Graph Embedding Extrapolate to Unseen Data: A Semantic Evidence ViewRen Li, Yanan Cao, Qiannan Zhu, Guanqun Bi 等AAAI 2022 · 被引用 103 次
- A Mean Field Analysis Of Deep ResNet And Beyond: Towards Provably Optimization Via Overparameterization From DepthYiping Lu, Chao Ma, Yulong Lu, Jianfeng Lu 等ICML 2020 · 被引用 85 次
- The Neural Covariance SDE: Shaped Infinite Depth-and-Width Networks at InitializationMufan Bill Li, Mihai Nica, Daniel M. RoyNeurIPS 2022 · 被引用 51 次
相关 Paper
- On the Role of Optimization in Double Descent: A Least Squares StudyIlja Kuzborskij, Csaba Szepesvári, Omar Rivasplata, Amal Rannen-Triki 等NeurIPS 2021 · 被引用 12 次
- The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of GeneralizationBen Adlam, Jeffrey PenningtonICML 2020 · 被引用 133 次
- On the Explicit Role of Initialization on the Convergence and Implicit Bias of Overparametrized Linear NetworksHancheng Min, Salma Tarmoun, René Vidal, Enrique MalladaICML 2021 · 被引用 53 次
- Double Trouble in Double Descent: Bias and Variance(s) in the Lazy RegimeStéphane d'Ascoli, Maria Refinetti, Giulio Biroli, Florent KrzakalaICML 2020 · 被引用 163 次
- Generalization Error of Generalized Linear Models in High DimensionsMelikasadat Emami, Mojtaba Sahraee-Ardakan, Parthe Pandit, Sundeep Rangan 等ICML 2020 · 被引用 40 次
