Hessian Eigenspectra of More Realistic Nonlinear Models
Zhenyu Liao, Michael W. Mahoney
摘要
Given an optimization problem, the Hessian matrix and its eigenspectrum can be used in many ways, ranging from designing more efficient second-order algorithms to performing model analysis and regression diagnostics. When nonlinear models and non-convex problems are considered, strong simplifying assumptions are often made to make Hessian spectral analysis more tractable. This leads to the question of how relevant the conclusions of such analyses are for more realistic nonlinear models. In this paper, we exploit deterministic equivalent techniques from random matrix theory to make a precise characterization of the Hessian eigenspectra for a broad family of nonlinear models, including models that generalize the classical generalized linear models, without relying on strong simplifying assumptions used previously. We show that, depending on the data properties, the nonlinear response model, and the loss function, the Hessian can have qualitatively different spectral behaviors: of bounded or unbounded support, with single-or multi-bulk, and with isolated eigenvalues on the left-or right-hand side of the bulk. By focusing on such a simple but nontrivial nonlinear model, our analysis takes a step forward to unveil the theoretical origin of many visually striking features observed in more complex machine learning models. • the (noisy) nonlinear factor model [22] , where y ∼ N (g(w T * x), σ 2 ) for some nonlinear linking function g : R → R and σ > 0; • the (noiseless) phase retrieval model [32] , with y = (w T * x) 2 , in which case we wish to reconstruct w * from its magnitude measurements; and • the single-layer NN model y = σ(w T * x), for some nonlinear activation function σ(t) such as the tanh-sigmoid σ(t) = tanh(t). For a given training set (x i , y i ) n i=1 of size n, the standard approach to obtain/recover the parameter w * ∈ R p is to solve the following optimization problem min w L(w) = min w 1 n n i=1 ℓ(y i , w T x i ), (3)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Why Transformers Need Adam: A Hessian PerspectiveYushun Zhang, Congliang Chen, Tian Ding, Ziniu Li 等NeurIPS 2024 · 被引用 149 次
- μPC: Scaling Predictive Coding to 100+ Layer NetworksFrancesco Innocenti, El Mehdi Achour, Christopher L. BuckleyNeurIPS 2025 · 被引用 19 次
- On the Overlooked Structure of Stochastic GradientsZeke Xie, Qian-Yuan Tang, Mingming Sun, Ping LiNeurIPS 2023 · 被引用 18 次
- Unifying Low Dimensional Spectra in Deep LearningConnall Garrod, Jonathan KeatingICML 2026 · 被引用 12 次
- Why Less is More (Sometimes): A Theory of Data CurationElvis Dohmatob, Mohammad Pezeshki, Reyhane Askari HemmatICLR 2026 · 被引用 11 次
它引用的顶会 Paper8
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERTSheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma 等AAAI 2020 · 被引用 656 次
- HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-PrecisionZhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney 等ICCV 2019 · 被引用 645 次
- ADAHESSIAN: An Adaptive Second Order Optimizer for Machine LearningZhewei Yao, Amir Gholami, Sheng Shen, Mustafa Mustafa 等AAAI 2021 · 被引用 358 次
- The Heavy-Tail Phenomenon in SGDMert Gürbüzbalaban, Umut Simsekli, Lingjiong ZhuICML 2021 · 被引用 165 次
- Multiplicative Noise and Heavy Tails in Stochastic OptimizationLiam Hodgkinson, Michael W. MahoneyICML 2021 · 被引用 90 次
相关 Paper
- Sparse Quantized Spectral ClusteringZhenyu Liao, Romain Couillet, Michael W. MahoneyICLR 2021 · 被引用 18 次
- Spectral Preconditioning for Gradient Methods on Graded Non-convex FunctionsNikita Doikov, Sebastian U. Stich, Martin JaggiICML 2024 · 被引用 10 次
- Optimal Iterative Sketching Methods with the Subsampled Randomized Hadamard TransformJonathan Lacotte, Sifan Liu, Edgar Dobriban, Mert PilanciNeurIPS 2020 · 被引用 15 次
- Analytic Characterization of the Hessian in Shallow ReLU Models: A Tale of SymmetryYossi Arjevani, Michael FieldNeurIPS 2020 · 被引用 22 次
- Analysis of Sensing Spectral for Signal Recovery under a Generalized Linear ModelJunjie Ma, Ji Xu, Arian MalekiNeurIPS 2021 · 被引用 10 次
