High-dimensional SGD aligns with emerging outlier eigenspaces
Gérard Ben Arous, Reza Gheissari, Jiaoyang Huang, Aukosh Jagannath
摘要
We rigorously study the relation between the training dynamics via stochastic gradient descent (SGD) and the spectra of empirical Hessian and gradient matrices. We prove that in two canonical classification tasks for multi-class high-dimensional mixtures and either 1 or 2-layer neural networks, both the SGD trajectory and emergent outlier eigenspaces of the Hessian and gradient matrices align with a common low-dimensional subspace. Moreover, in multi-layer settings this alignment occurs per layer, with the final layer's outlier eigenspace evolving over the course of training, and exhibiting rank deficiency when the SGD converges to sub-optimal classifiers. This establishes some of the rich predictions that have arisen from extensive numerical studies in the last decade about the spectra of Hessian and information matrices over the course of training in overparametrized networks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Linguistic Collapse: Neural Collapse in (Large) Language ModelsRobert Wu, Vardan PapyanNeurIPS 2024 · 被引用 45 次
- Unifying Low Dimensional Spectra in Deep LearningConnall Garrod, Jonathan KeatingICML 2026 · 被引用 12 次
- Efficient Sketches for Training Data Attribution and Studying the Loss LandscapeAndrea SchioppaNeurIPS 2024 · 被引用 11 次
- Understanding the Mechanisms of Fast Hyperparameter TransferNikhil Ghosh, Denny Wu, Alberto BiettiICLR 2026 · 被引用 8 次
- Deconstructing the Goldilocks Zone of Neural Network InitializationArtem Vysogorets, Anna Dawid, Julia KempeICML 2024 · 被引用 4 次
它引用的顶会 Paper14
- A Geometric Analysis of Neural Collapse with Unconstrained FeaturesZhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li 等NeurIPS 2021 · 被引用 303 次
- Neural Collapse Under MSE Loss: Proximity to and Dynamics on the Central PathX. Y. Han, Vardan Papyan, David L. DonohoICLR 2022 · 被引用 182 次
- Spectra of the Conjugate Kernel and Neural Tangent Kernel for linear-width neural networksZhou Fan, Zhichao WangNeurIPS 2020 · 被引用 101 次
- Dynamical mean-field theory for stochastic gradient descent in Gaussian mixture classificationFrancesca Mignacco, Florent Krzakala, Pierfrancesco Urbani, Lenka ZdeborováNeurIPS 2020 · 被引用 95 次
- High-dimensional limit theorems for SGD: Effective dynamics and critical scalingGérard Ben Arous, Reza Gheissari, Aukosh JagannathNeurIPS 2022 · 被引用 94 次
相关 Paper
- Does SGD really happen in tiny subspaces?Minhak Song, Kwangjun Ahn, Chulhee YunICLR 2025
- Spectral Evolution and Invariance in Linear-width Neural NetworksZhichao Wang, Andrew Engel, Anand D. Sarwate, Ioana Dumitriu 等NeurIPS 2023 · 被引用 33 次
- Chaotic Dynamics are Intrinsic to Neural Network Training with SGDLuis Herrmann, Maximilian Granz, Tim LandgrafNeurIPS 2022 · 被引用 15 次
- The asymptotic spectrum of the Hessian of DNN throughout trainingArthur Jacot, Franck Gabriel, Clément HonglerICLR 2020 · 被引用 39 次
- Trajectory Alignment: Understanding the Edge of Stability Phenomenon via Bifurcation TheoryMinhak Song, Chulhee YunNeurIPS 2023 · 被引用 26 次
