A Trace-restricted Kronecker-Factored Approximation to Natural Gradient
Kai-Xin Gao, Xiao-Lei Liu, Zheng-Hai Huang, Min Wang, Zidong Wang, Dachuan Xu, Fan Yu
Abstract
Second-order optimization methods have the ability to accelerate convergence by modifying the gradient through the curvature matrix. There have been many attempts to use second-order optimization methods for training deep neural networks. In this work, inspired by diagonal approximations and factored approximations such as Kronecker-factored Approximate Curvature (KFAC), we propose a new approximation to the Fisher information matrix (FIM) called Trace-restricted Kronecker-factored Approximate Curvature (TKFAC), which can hold the certain trace relationship between the exact and the approximate FIM. In TKFAC, we decompose each block of the approximate FIM as a Kronecker product of two smaller matrices and scaled by a coefficient related to trace. We theoretically analyze TKFAC's approximation error and give an upper bound of it. We also propose a new damping technique for TKFAC on convolutional neural networks to maintain the superiority of second-order optimization methods during training. Experiments show that our method has better performance compared with several state-of-the-art algorithms on some deep network architectures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bf427e9b-6f0a-497d-922a-02a11c7daecdCited by top-tier papers8
- NorMuon: Making Muon more efficient and scalableZichong Li, Liming Liu, Chen Liang, Weizhu Chen et al.ICML 2026 · 57 citations
- Second-Order Neural ODE OptimizerGuan-Horng Liu, Tianrong Chen, Evangelos A. TheodorouNeurIPS 2021 · 20 citations
- On the Parameterization of Second-Order Optimization Effective towards the Infinite WidthSatoki Ishikawa, Ryo KarakidaICLR 2024 · 10 citations
- Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank ExtensionWenbo Gong, Meyer Scetbon, Chao Ma, Edward MeedsICLR 2026 · 7 citations
- A Layer-Wise Natural Gradient Optimizer for Training Deep Neural NetworksXiaolei Liu, Shaoshuai Li, Kaixin Gao, Binfeng WangNeurIPS 2024 · 2 citations
Builds on3
- ADAHESSIAN: An Adaptive Second Order Optimizer for Machine LearningZhewei Yao, Amir Gholami, Sheng Shen, Mustafa Mustafa et al.AAAI 2021 · 358 citations
- Practical Quasi-Newton Methods for Training Deep Neural NetworksDonald Goldfarb, Yi Ren, Achraf BahamouNeurIPS 2020 · 130 citations
- Convolutional neural network training with distributed K-FACJ. Gregory Pauloski, Zhao Zhang, Lei Huang, Weijia Xu et al.SC 2020 · 26 citations
Related papers
- SKFAC: Training Neural Networks With Faster Kronecker-Factored Approximate CurvatureZedong Tang, Fenlong Jiang, Maoguo Gong, Hao Li et al.CVPR 2021
- Gradient Descent on Neurons and its Link to Approximate Second-order OptimizationFrederik BenzingICML 2022 · 31 citations
- Kronecker-Factored Approximate Curvature for Physics-Informed Neural NetworksFelix Dangel, Johannes Müller, Marius ZeinhoferNeurIPS 2024 · 31 citations
- THOR, Trace-based Hardware-driven Layer-Oriented Natural Gradient Descent ComputationMengyun Chen, Kai-Xin Gao, Xiaolei Liu, Zidong Wang et al.AAAI 2021 · 7 citations
- AdaFisher: Adaptive Second Order Optimization via Fisher InformationDamien Martins Gomes, Yanlei Zhang, Eugene Belilovsky, Guy Wolf et al.ICLR 2025
