Optimal criterion for feature learning of two-layer linear neural network in high dimensional interpolation regime
Keita Suzuki, Taiji Suzuki
摘要
Deep neural networks with feature learning have shown surprising generalization performance in high dimensional settings, but it has not been fully understood how and when they enjoy the benefit of feature learning. In this paper, we theoretically analyze the statistical properties of the benefit from feature learning in a two-layer linear neural network with multiple outputs in a high-dimensional setting. For that purpose, we propose a new criterion that allows feature lerning of a two-layer linear neural network in a high-dimensional setting. Interestingly, we can show that models with smaller values of the criterion generalize even in situations where normal ridge regression fails to generalize. This is because the proposed criterion contains a proper regularization for the feature mapping and acts as an upper bound on the predictive risk. As an important characterization of the criterion, the two-layer linear neural network that minimizes this criterion can achieve the optimal Bayes risk that is determined by the distribution of the true signals across the multiple outputs. To the best of our knowledge, this is the first study to specifically identify the conditions under which a model obtained by proper feature learning can outperform normal ridge regression in a high-dimensional multiple-output linear regression problem.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the RepresentationJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang 等NeurIPS 2022 · 被引用 173 次
- Benign Overfitting in Two-layer Convolutional Neural NetworksYuan Cao, Zixiang Chen, Misha Belkin, Quanquan GuNeurIPS 2022 · 被引用 121 次
- On the Implicit Bias of Initialization Shape: Beyond Infinitesimal Mirror DescentShahar Azulay, Edward Moroshko, Mor Shpigel Nacson, Blake E. Woodworth 等ICML 2021 · 被引用 85 次
- In Defense of Uniform Convergence: Generalization via Derandomization with an Application to Interpolating PredictorsJeffrey Negrea, Gintare Karolina Dziugaite, Daniel M. RoyICML 2020 · 被引用 66 次
相关 Paper
- Bayes-optimal Learning of Deep Random Networks of Extensive-widthHugo Cui, Florent Krzakala, Lenka ZdeborováICML 2023 · 被引用 49 次
- Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methodsTaiji Suzuki, Shunta AkiyamaICLR 2021 · 被引用 12 次
- High-Dimensional Analysis for Generalized Nonlinear Regression: From Asymptotics to AlgorithmJian Li, Yong Liu, Weiping WangAAAI 2024 · 被引用 4 次
- Generalization error in high-dimensional perceptrons: Approaching Bayes error with convex optimizationBenjamin Aubin, Florent Krzakala, Yue M. Lu, Lenka ZdeborováNeurIPS 2020 · 被引用 67 次
- A Theory of Non-Linear Feature Learning with One Gradient Step in Two-Layer Neural NetworksBehrad Moniri, Donghwan Lee, Hamed Hassani, Edgar DobribanICML 2024 · 被引用 38 次
