Lune

NeurIPS2022顶会

High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the Representation

Jimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang, Denny Wu, Greg Yang

2022年份
173被引次数
112顶会引用

摘要

We study the first gradient descent step on the first-layer parameters W\boldsymbol{W} in a two-layer neural network: f(x)=1Na⊤σ(W⊤x)f(\boldsymbol{x}) = \frac{1}{\sqrt{N}}\boldsymbol{a}^\top\sigma(\boldsymbol{W}^\top\boldsymbol{x}), where W∈Rd×N,a∈RN\boldsymbol{W}\in\mathbb{R}^{d\times N}, \boldsymbol{a}\in\mathbb{R}^{N} are randomly initialized, and the training objective is the empirical MSE loss: 1n∑i=1n(f(xi)−yi)2\frac{1}{n}\sum_{i=1}^n (f(\boldsymbol{x}_i)-y_i)^2. In the proportional asymptotic limit where n,d,N→∞n,d,N\to\infty at the same rate, and an idealized student-teacher setting, we show that the first gradient update contains a rank-1"spike", which results in an alignment between the first-layer weights and the linear component of the teacher model f∗f^*. To characterize the impact of this alignment, we compute the prediction risk of ridge regression on the conjugate kernel after one gradient step on W\boldsymbol{W} with learning rate η\eta, when f∗f^* is a single-index model. We consider two scalings of the first step learning rate η\eta. For small η\eta, we establish a Gaussian equivalence property for the trained feature map, and prove that the learned kernel improves upon the initial random features model, but cannot defeat the best linear model on the input. Whereas for sufficiently large η\eta, we prove that for certain f∗f^*, the same ridge estimator on trained features can go beyond this"linear regime"and outperform a wide range of random features and rotationally invariant kernels. Our results demonstrate that even one gradient step can lead to a considerable advantage over random features, and highlight the role of learning rate scaling in the initial phase of training.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper112

问问它们各自怎么用它

它引用的顶会 Paper24

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖