Lune

NeurIPS2025顶会

Learning quadratic neural networks in high dimensions: SGD dynamics and scaling laws

Gérard Ben Arous, Murat A. Erdogdu, Nuri Mert Vural, Denny Wu

2025年份
23被引次数
12顶会引用

摘要

We study the optimization and sample complexity of gradient-based training of a two-layer neural network with quadratic activation function in the high-dimensional regime, where the data is generated as f∗(x)∝∑j=1rλjσ(⟨θj,x⟩),x∼N(0,Id)f_*(\boldsymbol{x}) \propto \sum_{j=1}^{r}\lambda_j \sigma\left(\langle \boldsymbol{\theta_j}, \boldsymbol{x}\rangle\right), \boldsymbol{x} \sim N(0,\boldsymbol{I}_d), σ\sigma is the 2nd Hermite polynomial, and {θj}j=1r⊂Rd\lbrace\boldsymbol{\theta}_j \rbrace_{j=1}^{r} \subset \mathbb{R}^d are orthonormal signal directions. We consider the extensive-width regime r≍dβr \asymp d^\beta for β∈[0,1)\beta \in [0, 1), and assume a power-law decay on the (non-negative) second-layer coefficients λj≍j−α\lambda_j\asymp j^{-\alpha} for α≥0\alpha \geq 0. We present a sharp analysis of the SGD dynamics in the feature learning regime, for both the population limit and the finite-sample (online) discretization, and derive scaling laws for the prediction risk that highlight the power-law dependencies on the optimization time, sample size, and model width. Our analysis combines a precise characterization of the associated matrix Riccati differential equation with novel matrix monotonicity arguments to establish convergence guarantees for the infinite-dimensional effective dynamics.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper12

问问它们各自怎么用它

它引用的顶会 Paper21

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖