Gradient Descent on Neurons and its Link to Approximate Second-order Optimization
Frederik Benzing
Abstract
Second-order optimizers are thought to hold the potential to speed up neural network training, but due to the enormous size of the curvature matrix, they typically require approximations to be computationally tractable. The most successful family of approximations are Kronecker-Factored, block-diagonal curvature estimates (KFAC). Here, we combine tools from prior work to evaluate exact second-order updates with careful ablations to establish a surprising result: Due to its approximations, KFAC is not closely related to second-order updates, and in particular, it significantly outperforms true second-order updates. This challenges widely held believes and immediately raises the question why KFAC performs so well. Towards answering this question we present evidence strongly suggesting that KFAC approximates a first-order algorithm, which performs gradient descent on neurons rather than weights. Finally, we show that this optimizer often improves over KFAC in terms of computational cost and data-efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5b9c51fd-57c5-45c6-b51c-7236e63c1fa1Cited by top-tier papers17
- Kronecker-Factored Approximate Curvature for Modern Neural Network ArchitecturesRuna Eschenhagen, Alexander Immer, Richard E. Turner, Frank Schneider et al.NeurIPS 2023 · 62 citations
- FedLPA: One-shot Federated Learning with Layer-Wise Posterior AggregationXiang Liu, Liangxi Liu, Feiyang Ye, Yunheng Shen et al.NeurIPS 2024 · 34 citations
- Kronecker-Factored Approximate Curvature for Physics-Informed Neural NetworksFelix Dangel, Johannes Müller, Marius ZeinhoferNeurIPS 2024 · 31 citations
- An Improved Empirical Fisher Approximation for Natural Gradient DescentXiaodong Wu, Wenyi Yu, Chao Zhang, Philip C. WoodlandNeurIPS 2024 · 27 citations
- Purifying Shampoo: Investigating Shampoo's Heuristics by Decomposing its PreconditionerRuna Eschenhagen, Aaron Defazio, Tsung-Hsien Lee, Richard E. Turner et al.NeurIPS 2025 · 24 citations
Builds on6
- Laplace Redux - Effortless Bayesian Deep LearningErik A. Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen et al.NeurIPS 2021 · 508 citations
- Scalable Marginal Likelihood Estimation for Model Selection in Deep LearningAlexander Immer, Matthias Bauer, Vincent Fortuin, Gunnar Rätsch et al.ICML 2021 · 130 citations
- Practical Quasi-Newton Methods for Training Deep Neural NetworksDonald Goldfarb, Yi Ren, Achraf BahamouNeurIPS 2020 · 130 citations
- A Theoretical Framework for Target PropagationAlexander Meulemans, Francesco S. Carzaniga, Johan A. K. Suykens, João Sacramento et al.NeurIPS 2020 · 110 citations
- Global inducing point variational posteriors for Bayesian neural networks and deep Gaussian processesSebastian W. Ober, Laurence AitchisonICML 2021 · 65 citations
Related papers
- SKFAC: Training Neural Networks With Faster Kronecker-Factored Approximate CurvatureZedong Tang, Fenlong Jiang, Maoguo Gong, Hao Li et al.CVPR 2021
- A Trace-restricted Kronecker-Factored Approximation to Natural GradientKai-Xin Gao, Xiao-Lei Liu, Zheng-Hai Huang, Min Wang et al.AAAI 2021 · 13 citations
- Convolutional neural network training with distributed K-FACJ. Gregory Pauloski, Zhao Zhang, Lei Huang, Weijia Xu et al.SC 2020 · 26 citations
- KAISA: an adaptive second-order optimizer framework for deep neural networksJ. Gregory Pauloski, Qi Huang, Lei Huang, Shivaram Venkataraman et al.SC 2021 · 14 citations
- Structured Inverse-Free Natural Gradient Descent: Memory-Efficient & Numerically-Stable KFACWu Lin, Felix Dangel, Runa Eschenhagen, Kirill Neklyudov et al.ICML 2024 · 7 citations
