Lune

NeurIPS2023顶会

On Single-Index Models beyond Gaussian Data

Aaron Zweig, Loucas Pillaud-Vivien, Joan Bruna

2023年份
17被引次数
10顶会引用

摘要

Sparse high-dimensional functions have arisen as a rich framework to study the behavior of gradient-descent methods using shallow neural networks, showcasing their ability to perform feature learning beyond linear models. Amongst those functions, the simplest are single-index models f(x)=ϕ(x⋅θ∗)f(x) = \phi( x \cdot \theta^*), where the labels are generated by an arbitrary non-linear scalar link function ϕ\phi applied to an unknown one-dimensional projection θ∗\theta^* of the input data. By focusing on Gaussian data, several recent works have built a remarkable picture, where the so-called information exponent (related to the regularity of the link function) controls the required sample complexity. In essence, these tools exploit the stability and spherical symmetry of Gaussian distributions. In this work, building from the framework of , we explore extensions of this picture beyond the Gaussian setting, where both stability or symmetry might be violated. Focusing on the planted setting where ϕ\phi is known, our main results establish that Stochastic Gradient Descent can efficiently recover the unknown direction θ∗\theta^* in the high-dimensional regime, under assumptions that extend previous works .

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper10

问问它们各自怎么用它

它引用的顶会 Paper4

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖