Analytic Insights into Structure and Rank of Neural Network Hessian Maps
Sidak Pal Singh, Gregor Bachmann, Thomas Hofmann
Abstract
The Hessian of a neural network captures parameter interactions through secondorder derivatives of the loss. It is a fundamental object of study, closely tied to various problems in deep learning, including model design, optimization, and generalization. Most prior work has been empirical, typically focusing on lowrank approximations and heuristics that are blind to the network structure. In contrast, we develop theoretical tools to analyze the range of the Hessian map, providing us with a precise understanding of its rank deficiency as well as the structural reasons behind it. This yields exact formulas and tight upper bounds for the Hessian rank of deep linear networks, allowing for an elegant interpretation in terms of rank deficiency. Moreover, we demonstrate that our bounds remain faithful as an estimate of the numerical Hessian rank, for a larger class of models such as rectified and hyperbolic tangent networks. Further, we also investigate the implications of model architecture (e.g. width, depth, bias) on the rank deficiency. Overall, our work provides novel insights into the source and extent of redundancy in overparameterized networks. * Detailed list of contributions are: Sidak first discovered that the Hessian rank formula, in an early form, holds experimentally to high fidelity, thus kick-starting the project. Sidak came up with the proof technique and proved Theorem 3, Theorem 5, Theorem 9, Theorem 12. Sidak wrote essentially the entire paper and noted the rank-deficiency interpretation. Gregor proved Lemma 8, assisted in a part of Theorem 3, and empirically observed the eventual formula for the Hessian rank. Gregor essentially ran all the experiments for the final submission and made the corresponding figures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f630b10b-bbaf-4ba4-aa9d-c06ab086b450Cited by top-tier papers24
- ReLoRA: High-Rank Training Through Low-Rank UpdatesVladislav Lialin, Sherin Muckatira, Namrata Shivagunde, Anna RumshiskyICLR 2024 · 214 citations
- Signal Propagation in Transformers: Theoretical Perspectives and the Role of Rank CollapseLorenzo Noci, Sotiris Anagnostidis, Luca Biggio, Antonio Orvieto et al.NeurIPS 2022 · 161 citations
- Low-rank lottery tickets: finding efficient low-rank neural networks via matrix differential equationsSteffen Schotthöfer, Emanuele Zangrando, Jonas Kusch, Gianluca Ceruti et al.NeurIPS 2022 · 66 citations
- Super Consistency of Neural Network Landscapes and Learning Rate TransferLorenzo Noci, Alexandru Meterez, Thomas Hofmann, Antonio OrvietoNeurIPS 2024 · 25 citations
- Robust low-rank training via approximate orthonormal constraintsDayana Savostianova, Emanuele Zangrando, Gianluca Ceruti, Francesco TudiscoNeurIPS 2023 · 24 citations
Builds on4
- WoodFisher: Efficient Second-Order Approximation for Neural Network CompressionSidak Pal Singh, Dan AlistarhNeurIPS 2020 · 217 citations
- Practical Quasi-Newton Methods for Training Deep Neural NetworksDonald Goldfarb, Yi Ren, Achraf BahamouNeurIPS 2020 · 130 citations
- The asymptotic spectrum of the Hessian of DNN throughout trainingArthur Jacot, Franck Gabriel, Clément HonglerICLR 2020 · 39 citations
- Adaptive Newton Sketch: Linear-time Optimization with Quadratic Convergence and Effective Hessian DimensionalityJonathan Lacotte, Yifei Wang, Mert PilanciICML 2021 · 18 citations
Related papers
- The Hessian perspective into the Nature of Convolutional Neural NetworksSidak Pal Singh, Thomas Hofmann, Bernhard SchölkopfICML 2023 · 12 citations
- Rank Diminishing in Deep Neural NetworksRuili Feng, Kecheng Zheng, Yukun Huang, Deli Zhao et al.NeurIPS 2022 · 64 citations
- Unifying Low Dimensional Spectra in Deep LearningConnall Garrod, Jonathan KeatingICML 2026 · 12 citations
- What can linearized neural networks actually say about generalization?Guillermo Ortiz-Jiménez, Seyed-Mohsen Moosavi-Dezfooli, Pascal FrossardNeurIPS 2021 · 62 citations
- Phenomenology of Double Descent in Finite-Width Neural NetworksSidak Pal Singh, Aurélien Lucchi, Thomas Hofmann, Bernhard SchölkopfICLR 2022 · 12 citations
