Lune

NeurIPS2023顶会

Smoothing the Landscape Boosts the Signal for SGD: Optimal Sample Complexity for Learning Single Index Models

Alex Damian, Eshaan Nichani, Rong Ge, Jason D. Lee

2023年份
67被引次数
38顶会引用

摘要

We focus on the task of learning a single index model σ(w⋆⋅x)\sigma(w^\star \cdot x) with respect to the isotropic Gaussian distribution in dd dimensions. Prior work has shown that the sample complexity of learning w⋆w^\star is governed by the information exponent k⋆k^\star of the link function σ\sigma, which is defined as the index of the first nonzero Hermite coefficient of σ\sigma. Ben Arous et al. (2021) showed that n≳dk⋆−1n \gtrsim d^{k^\star-1} samples suffice for learning w⋆w^\star and that this is tight for online SGD. However, the CSQ lower bound for gradient based methods only shows that n≳dk⋆/2n \gtrsim d^{k^\star/2} samples are necessary. In this work, we close the gap between the upper and lower bounds by showing that online SGD on a smoothed loss learns w⋆w^\star with n≳dk⋆/2n \gtrsim d^{k^\star/2} samples. We also draw connections to statistical analyses of tensor PCA and to the implicit regularization effects of minibatch SGD on empirical losses.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper38

问问它们各自怎么用它

它引用的顶会 Paper6

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖