Lune

NeurIPS2023Top-tier venue

Learning in the Presence of Low-dimensional Structure: A Spiked Random Matrix Perspective

Jimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang, Denny Wu

2023Year
47Citations
26Top-tier citations

Abstract

We consider the problem of learning a single-index target function f * : R d → R under the spiked covariance data: where the link function σ * : R → R is a degree-p polynomial with information exponent k (defined as the lowest degree in the Hermite expansion of σ * ), and it depends on the projection of input x onto the spike (signal) direction µ ∈ R d . In the proportional asymptotic limit where the number of training examples n and the dimensionality d jointly diverge: n, d → ∞, n/d → ψ ∈ (0, ∞), we ask the following question: how large should the spike magnitude θ be, in order for (i) kernel methods, (ii) neural networks optimized by gradient descent, to learn f * ? We show that for kernel ridge regression, β ≥ 1 -1 p is both sufficient and necessary. Whereas for two-layer neural networks trained with gradient descent, β > 1 -1 k suffices. Our results demonstrate that both kernel methods and neural networks benefit from low-dimensional structures in the data. Further, since k ≤ p by definition, neural networks can adapt to such structures more effectively.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext d943c4fd-c8c6-48b0-8966-6eccfe50a7d2

Cited by top-tier papers26

Ask how each one uses it

Builds on8

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines