A Deeper Look at the Hessian Eigenspectrum of Deep Neural Networks and its Applications to Regularization
Adepu Ravi Sankar, Yash Khasbage, Rahul Vigneswaran, Vineeth N. Balasubramanian
摘要
Loss landscape analysis is extremely useful for a deeper understanding of the generalization ability of deep neural network models. In this work, we propose a layerwise loss landscape analysis where the loss surface at every layer is studied independently and also on how each correlates to the overall loss surface. We study the layerwise loss landscape by studying the eigenspectra of the Hessian at each layer. In particular, our results show that the layerwise Hessian geometry is largely similar to the entire Hessian. We also report an interesting phenomenon where the Hessian eigenspectrum of middle layers of the deep neural network are observed to most similar to the overall Hessian eigenspectrum. We also show that the maximum eigenvalue and the trace of the Hessian (both full network and layerwise) reduce as training of the network progresses. We leverage on these observations to propose a new regularizer based on the trace of the layerwise Hessian. Penalizing the trace of the Hessian at every layer indirectly forces Stochastic Gradient Descent to converge to flatter minima, which are shown to have better generalization performance. In particular, we show that such a layerwise regularizer can be leveraged to penalize the middlemost layers alone, which yields promising results. Our empirical studies on well-known deep nets across datasets support the claims of this work.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-trainingHong Liu, Zhiyuan Li, David Leo Wright Hall, Percy Liang 等ICLR 2024 · 被引用 264 次
- Why Transformers Need Adam: A Hessian PerspectiveYushun Zhang, Congliang Chen, Tian Ding, Ziniu Li 等NeurIPS 2024 · 被引用 149 次
- When Do Flat Minima Optimizers Work?Jean Kaddour, Linqing Liu, Ricardo Silva, Matt J. KusnerNeurIPS 2022 · 被引用 102 次
- Gradient Alignment in Physics-informed Neural Networks: A Second-Order Optimization PerspectiveSifan Wang, Ananyae Kumar Bhartari, Bowen Li, Paris PerdikarisNeurIPS 2025 · 被引用 100 次
- Hessian Eigenspectra of More Realistic Nonlinear ModelsZhenyu Liao, Michael W. MahoneyNeurIPS 2021 · 被引用 45 次
相关 Paper
- How Sharpness-Aware Minimization Minimizes Sharpness?Kaiyue Wen, Tengyu Ma, Zhiyuan LiICLR 2023 · 被引用 3 次
- CR-SAM: Curvature Regularized Sharpness-Aware MinimizationTao Wu, Tie Luo, Donald C. Wunsch IIAAAI 2024 · 被引用 15 次
- Investigating the Overlooked Hessian Structure: From CNNs to LLMsQian-Yuan Tang, Yufei Gu, Yunfeng Cai, Mingming Sun 等ICML 2025
- Boosting Adversarial Transferability via Negative Hessian Trace RegularizationYunfei Long, Zilin Tian, Liguo Zhang, Huosheng XuICCV 2025 · 被引用 1 次
- Phase diagram of early training dynamics in deep neural networks: effect of the learning rate, depth, and widthDayal Singh Kalra, Maissam BarkeshliNeurIPS 2023 · 被引用 21 次
