Hessian Eigenspectra of More Realistic Nonlinear Models
Zhenyu Liao, Michael W. Mahoney
Abstract
Given an optimization problem, the Hessian matrix and its eigenspectrum can be used in many ways, ranging from designing more efficient second-order algorithms to performing model analysis and regression diagnostics. When nonlinear models and non-convex problems are considered, strong simplifying assumptions are often made to make Hessian spectral analysis more tractable. This leads to the question of how relevant the conclusions of such analyses are for more realistic nonlinear models. In this paper, we exploit deterministic equivalent techniques from random matrix theory to make a precise characterization of the Hessian eigenspectra for a broad family of nonlinear models, including models that generalize the classical generalized linear models, without relying on strong simplifying assumptions used previously. We show that, depending on the data properties, the nonlinear response model, and the loss function, the Hessian can have qualitatively different spectral behaviors: of bounded or unbounded support, with single-or multi-bulk, and with isolated eigenvalues on the left-or right-hand side of the bulk. By focusing on such a simple but nontrivial nonlinear model, our analysis takes a step forward to unveil the theoretical origin of many visually striking features observed in more complex machine learning models. • the (noisy) nonlinear factor model [22] , where y ∼ N (g(w T * x), σ 2 ) for some nonlinear linking function g : R → R and σ > 0; • the (noiseless) phase retrieval model [32] , with y = (w T * x) 2 , in which case we wish to reconstruct w * from its magnitude measurements; and • the single-layer NN model y = σ(w T * x), for some nonlinear activation function σ(t) such as the tanh-sigmoid σ(t) = tanh(t). For a given training set (x i , y i ) n i=1 of size n, the standard approach to obtain/recover the parameter w * ∈ R p is to solve the following optimization problem min w L(w) = min w 1 n n i=1 ℓ(y i , w T x i ), (3)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers19
- Why Transformers Need Adam: A Hessian PerspectiveYushun Zhang, Congliang Chen, Tian Ding, Ziniu Li et al.NeurIPS 2024 · 149 citations
- μPC: Scaling Predictive Coding to 100+ Layer NetworksFrancesco Innocenti, El Mehdi Achour, Christopher L. BuckleyNeurIPS 2025 · 19 citations
- On the Overlooked Structure of Stochastic GradientsZeke Xie, Qian-Yuan Tang, Mingming Sun, Ping LiNeurIPS 2023 · 18 citations
- Unifying Low Dimensional Spectra in Deep LearningConnall Garrod, Jonathan KeatingICML 2026 · 12 citations
- Why Less is More (Sometimes): A Theory of Data CurationElvis Dohmatob, Mohammad Pezeshki, Reyhane Askari HemmatICLR 2026 · 11 citations
Builds on8
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERTSheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma et al.AAAI 2020 · 656 citations
- HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-PrecisionZhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney et al.ICCV 2019 · 645 citations
- ADAHESSIAN: An Adaptive Second Order Optimizer for Machine LearningZhewei Yao, Amir Gholami, Sheng Shen, Mustafa Mustafa et al.AAAI 2021 · 358 citations
- The Heavy-Tail Phenomenon in SGDMert Gürbüzbalaban, Umut Simsekli, Lingjiong ZhuICML 2021 · 165 citations
- Multiplicative Noise and Heavy Tails in Stochastic OptimizationLiam Hodgkinson, Michael W. MahoneyICML 2021 · 90 citations
Related papers
- Sparse Quantized Spectral ClusteringZhenyu Liao, Romain Couillet, Michael W. MahoneyICLR 2021 · 18 citations
- Spectral Preconditioning for Gradient Methods on Graded Non-convex FunctionsNikita Doikov, Sebastian U. Stich, Martin JaggiICML 2024 · 10 citations
- Optimal Iterative Sketching Methods with the Subsampled Randomized Hadamard TransformJonathan Lacotte, Sifan Liu, Edgar Dobriban, Mert PilanciNeurIPS 2020 · 15 citations
- Analytic Characterization of the Hessian in Shallow ReLU Models: A Tale of SymmetryYossi Arjevani, Michael FieldNeurIPS 2020 · 22 citations
- Analysis of Sensing Spectral for Signal Recovery under a Generalized Linear ModelJunjie Ma, Ji Xu, Arian MalekiNeurIPS 2021 · 10 citations
