High-dimensional SGD aligns with emerging outlier eigenspaces
Gérard Ben Arous, Reza Gheissari, Jiaoyang Huang, Aukosh Jagannath
Abstract
We rigorously study the relation between the training dynamics via stochastic gradient descent (SGD) and the spectra of empirical Hessian and gradient matrices. We prove that in two canonical classification tasks for multi-class high-dimensional mixtures and either 1 or 2-layer neural networks, both the SGD trajectory and emergent outlier eigenspaces of the Hessian and gradient matrices align with a common low-dimensional subspace. Moreover, in multi-layer settings this alignment occurs per layer, with the final layer's outlier eigenspace evolving over the course of training, and exhibiting rank deficiency when the SGD converges to sub-optimal classifiers. This establishes some of the rich predictions that have arisen from extensive numerical studies in the last decade about the spectra of Hessian and information matrices over the course of training in overparametrized networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Linguistic Collapse: Neural Collapse in (Large) Language ModelsRobert Wu, Vardan PapyanNeurIPS 2024 · 45 citations
- Unifying Low Dimensional Spectra in Deep LearningConnall Garrod, Jonathan KeatingICML 2026 · 12 citations
- Efficient Sketches for Training Data Attribution and Studying the Loss LandscapeAndrea SchioppaNeurIPS 2024 · 11 citations
- Understanding the Mechanisms of Fast Hyperparameter TransferNikhil Ghosh, Denny Wu, Alberto BiettiICLR 2026 · 8 citations
- Deconstructing the Goldilocks Zone of Neural Network InitializationArtem Vysogorets, Anna Dawid, Julia KempeICML 2024 · 4 citations
Builds on14
- A Geometric Analysis of Neural Collapse with Unconstrained FeaturesZhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li et al.NeurIPS 2021 · 303 citations
- Neural Collapse Under MSE Loss: Proximity to and Dynamics on the Central PathX. Y. Han, Vardan Papyan, David L. DonohoICLR 2022 · 182 citations
- Spectra of the Conjugate Kernel and Neural Tangent Kernel for linear-width neural networksZhou Fan, Zhichao WangNeurIPS 2020 · 101 citations
- Dynamical mean-field theory for stochastic gradient descent in Gaussian mixture classificationFrancesca Mignacco, Florent Krzakala, Pierfrancesco Urbani, Lenka ZdeborováNeurIPS 2020 · 95 citations
- High-dimensional limit theorems for SGD: Effective dynamics and critical scalingGérard Ben Arous, Reza Gheissari, Aukosh JagannathNeurIPS 2022 · 94 citations
Related papers
- Does SGD really happen in tiny subspaces?Minhak Song, Kwangjun Ahn, Chulhee YunICLR 2025
- Spectral Evolution and Invariance in Linear-width Neural NetworksZhichao Wang, Andrew Engel, Anand D. Sarwate, Ioana Dumitriu et al.NeurIPS 2023 · 33 citations
- Chaotic Dynamics are Intrinsic to Neural Network Training with SGDLuis Herrmann, Maximilian Granz, Tim LandgrafNeurIPS 2022 · 15 citations
- The asymptotic spectrum of the Hessian of DNN throughout trainingArthur Jacot, Franck Gabriel, Clément HonglerICLR 2020 · 39 citations
- Trajectory Alignment: Understanding the Edge of Stability Phenomenon via Bifurcation TheoryMinhak Song, Chulhee YunNeurIPS 2023 · 26 citations
