Gradient Descent on Neurons and its Link to Approximate Second-order Optimization
Frederik Benzing
摘要
Second-order optimizers are thought to hold the potential to speed up neural network training, but due to the enormous size of the curvature matrix, they typically require approximations to be computationally tractable. The most successful family of approximations are Kronecker-Factored, block-diagonal curvature estimates (KFAC). Here, we combine tools from prior work to evaluate exact second-order updates with careful ablations to establish a surprising result: Due to its approximations, KFAC is not closely related to second-order updates, and in particular, it significantly outperforms true second-order updates. This challenges widely held believes and immediately raises the question why KFAC performs so well. Towards answering this question we present evidence strongly suggesting that KFAC approximates a first-order algorithm, which performs gradient descent on neurons rather than weights. Finally, we show that this optimizer often improves over KFAC in terms of computational cost and data-efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Kronecker-Factored Approximate Curvature for Modern Neural Network ArchitecturesRuna Eschenhagen, Alexander Immer, Richard E. Turner, Frank Schneider 等NeurIPS 2023 · 被引用 62 次
- FedLPA: One-shot Federated Learning with Layer-Wise Posterior AggregationXiang Liu, Liangxi Liu, Feiyang Ye, Yunheng Shen 等NeurIPS 2024 · 被引用 34 次
- Kronecker-Factored Approximate Curvature for Physics-Informed Neural NetworksFelix Dangel, Johannes Müller, Marius ZeinhoferNeurIPS 2024 · 被引用 31 次
- An Improved Empirical Fisher Approximation for Natural Gradient DescentXiaodong Wu, Wenyi Yu, Chao Zhang, Philip C. WoodlandNeurIPS 2024 · 被引用 27 次
- Purifying Shampoo: Investigating Shampoo's Heuristics by Decomposing its PreconditionerRuna Eschenhagen, Aaron Defazio, Tsung-Hsien Lee, Richard E. Turner 等NeurIPS 2025 · 被引用 24 次
它引用的顶会 Paper6
- Laplace Redux - Effortless Bayesian Deep LearningErik A. Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen 等NeurIPS 2021 · 被引用 508 次
- Scalable Marginal Likelihood Estimation for Model Selection in Deep LearningAlexander Immer, Matthias Bauer, Vincent Fortuin, Gunnar Rätsch 等ICML 2021 · 被引用 130 次
- Practical Quasi-Newton Methods for Training Deep Neural NetworksDonald Goldfarb, Yi Ren, Achraf BahamouNeurIPS 2020 · 被引用 130 次
- A Theoretical Framework for Target PropagationAlexander Meulemans, Francesco S. Carzaniga, Johan A. K. Suykens, João Sacramento 等NeurIPS 2020 · 被引用 110 次
- Global inducing point variational posteriors for Bayesian neural networks and deep Gaussian processesSebastian W. Ober, Laurence AitchisonICML 2021 · 被引用 65 次
相关 Paper
- SKFAC: Training Neural Networks With Faster Kronecker-Factored Approximate CurvatureZedong Tang, Fenlong Jiang, Maoguo Gong, Hao Li 等CVPR 2021
- A Trace-restricted Kronecker-Factored Approximation to Natural GradientKai-Xin Gao, Xiao-Lei Liu, Zheng-Hai Huang, Min Wang 等AAAI 2021 · 被引用 13 次
- Convolutional neural network training with distributed K-FACJ. Gregory Pauloski, Zhao Zhang, Lei Huang, Weijia Xu 等SC 2020 · 被引用 26 次
- KAISA: an adaptive second-order optimizer framework for deep neural networksJ. Gregory Pauloski, Qi Huang, Lei Huang, Shivaram Venkataraman 等SC 2021 · 被引用 14 次
- Structured Inverse-Free Natural Gradient Descent: Memory-Efficient & Numerically-Stable KFACWu Lin, Felix Dangel, Runa Eschenhagen, Kirill Neklyudov 等ICML 2024 · 被引用 7 次
