The asymptotic spectrum of the Hessian of DNN throughout training
Arthur Jacot, Franck Gabriel, Clément Hongler
Abstract
The dynamics of DNNs during gradient descent is described by the so-called Neural Tangent Kernel (NTK). In this article, we show that the NTK allows one to gain precise insight into the Hessian of the cost of DNNs. When the NTK is fixed during training, we obtain a full characterization of the asymptotics of the spectrum of the Hessian, at initialization and during training. In the so-called mean-field limit, where the NTK is not fixed during training, we describe the first two moments of the Hessian at initialization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 15aa598d-192e-4e04-8c50-19a08c1a19e4Cited by top-tier papers17
- Spectra of the Conjugate Kernel and Neural Tangent Kernel for linear-width neural networksZhou Fan, Zhichao WangNeurIPS 2020 · 101 citations
- Analytic Insights into Structure and Rank of Neural Network Hessian MapsSidak Pal Singh, Gregor Bachmann, Thomas HofmannNeurIPS 2021 · 60 citations
- Hessian Eigenspectra of More Realistic Nonlinear ModelsZhenyu Liao, Michael W. MahoneyNeurIPS 2021 · 45 citations
- Second-order regression models exhibit progressive sharpening to the edge of stabilityAtish Agarwala, Fabian Pedregosa, Jeffrey PenningtonICML 2023 · 37 citations
- Global Convergence of MAML and Theory-Inspired Neural Architecture Search for Few-Shot LearningHaoxiang Wang, Yite Wang, Ruoyu Sun, Bo LiCVPR 2022 · 37 citations
Builds on1
Related papers
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent KernelStanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani et al.NeurIPS 2020 · 255 citations
- On the linearity of large non-linear models: when and why the tangent kernel is constantChaoyue Liu, Libin Zhu, Mikhail BelkinNeurIPS 2020 · 183 citations
- The Challenges of the Nonlinear Regime for Physics-Informed Neural NetworksAndrea Bonfanti, Giuseppe Bruno, Cristina CiprianiNeurIPS 2024 · 41 citations
- The Surprising Simplicity of the Early-Time Learning Dynamics of Neural NetworksWei Hu, Lechao Xiao, Ben Adlam, Jeffrey PenningtonNeurIPS 2020 · 77 citations
- Characterizing the spectrum of the NTK via a power series expansionMichael Murray, Hui Jin, Benjamin Bowman, Guido MontúfarICLR 2023 · 2 citations
