Low Rank Gradients and Where to Find Them
Rishi Sonthalia, Michael Murray, Guido F. Montúfar
Abstract
This paper investigates low-rank structure in the gradients of the training loss for two-layer neural networks while relaxing the usual isotropy assumptions on the training data and parameters. We consider a spiked data model in which the bulk can be anisotropic and ill-conditioned, we do not require independent data and weight matrices and we also analyze both the mean-field and neural-tangent-kernel scalings. We show that the gradient with respect to the input weights is approximately low rank and is dominated by two rank-one terms: one aligned with the bulk data-residue , and another aligned with the rank one spike in the input data. We characterize how properties of the training data, the scaling regime and the activation function govern the balance between these two components. Additionally, we also demonstrate that standard regularizers, such as weight decay, input noise and Jacobian penalties, also selectively modulate these components. Experiments on synthetic and real data corroborate our theoretical predictions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on13
- On the linearity of large non-linear models: when and why the tangent kernel is constantChaoyue Liu, Libin Zhu, Mikhail BelkinNeurIPS 2020 · 183 citations
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the RepresentationJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang et al.NeurIPS 2022 · 173 citations
- Spectra of the Conjugate Kernel and Neural Tangent Kernel for linear-width neural networksZhou Fan, Zhichao WangNeurIPS 2020 · 101 citations
- The inductive bias of ReLU networks on orthogonally separable dataMary Phuong, Christoph H. LampertICLR 2021 · 53 citations
- Learning in the Presence of Low-dimensional Structure: A Spiked Random Matrix PerspectiveJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang et al.NeurIPS 2023 · 47 citations
Related papers
- A Generalized Neural Tangent Kernel Analysis for Two-layer Neural NetworksZixiang Chen, Yuan Cao, Quanquan Gu, Tong ZhangNeurIPS 2020 · 82 citations
- Gradient-Based Feature Learning under Structured DataAlireza Mousavi-Hosseini, Denny Wu, Taiji Suzuki, Murat A. ErdogduNeurIPS 2023 · 36 citations
- Spectral Evolution and Invariance in Linear-width Neural NetworksZhichao Wang, Andrew Engel, Anand D. Sarwate, Ioana Dumitriu et al.NeurIPS 2023 · 33 citations
- Wide Neural Networks Trained with Weight Decay Provably Exhibit Neural CollapseArthur Jacot, Peter Súkeník, Zihan Wang, Marco MondelliICLR 2025
- Implicit Bias in Leaky ReLU Networks Trained on High-Dimensional DataSpencer Frei, Gal Vardi, Peter L. Bartlett, Nathan Srebro et al.ICLR 2023 · 5 citations
