SKFAC: Training Neural Networks With Faster Kronecker-Factored Approximate Curvature
Zedong Tang, Fenlong Jiang, Maoguo Gong, Hao Li, Yue Wu, Fan Yu, Zidong Wang, Min Wang
摘要
The bottleneck of computation burden limits the widespread use of the 2nd order optimization algorithms for training deep neural networks. In this paper, we present a computationally efficient approximation for natural gradient descent, named Swift Kronecker-Factored Approximate Curvature (SKFAC), which combines Kronecker factorization and a fast low-rank matrix inversion technique. Our research aims at both fully connected and convolutional layers. For the fully connected layers, by utilizing the low-rank property of Kronecker factors of Fisher information matrix, our method only requires inverting a small matrix to approximate the curvature with desirable accuracy. For convolutional layers, we propose a way with two strategies to save computational efforts without affecting the empirical performance by reducing across the spatial dimension or receptive fields of feature maps. Specifically, we propose two effective dimension reduction methods for this purpose: Spatial Subsampling and Reduce Sum. Experimental results of training several deep neural networks on Cifar-10 and ImageNet-1k datasets demonstrate that SKFAC can capture the main curvature and yield comparative performance to K-FAC. The proposed method bridges the wall-clock time gap between the 1st and 2nd order algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Kronecker-Factored Approximate Curvature for Modern Neural Network ArchitecturesRuna Eschenhagen, Alexander Immer, Richard E. Turner, Frank Schneider 等NeurIPS 2023 · 被引用 62 次
- Amortized Proximal OptimizationJuhan Bae, Paul Vicol, Jeff Z. HaoChen, Roger B. GrosseNeurIPS 2022 · 被引用 15 次
- AdaFisher: Adaptive Second Order Optimization via Fisher InformationDamien Martins Gomes, Yanlei Zhang, Eugene Belilovsky, Guy Wolf 等ICLR 2025
- Influence Functions for Scalable Data Attribution in Diffusion ModelsBruno Kacper Mlodozeniec, Runa Eschenhagen, Juhan Bae, Alexander Immer 等ICLR 2025
相关 Paper
- A Trace-restricted Kronecker-Factored Approximation to Natural GradientKai-Xin Gao, Xiao-Lei Liu, Zheng-Hai Huang, Min Wang 等AAAI 2021 · 被引用 13 次
- Convolutional neural network training with distributed K-FACJ. Gregory Pauloski, Zhao Zhang, Lei Huang, Weijia Xu 等SC 2020 · 被引用 26 次
- Gradient Descent on Neurons and its Link to Approximate Second-order OptimizationFrederik BenzingICML 2022 · 被引用 31 次
- HyLo: A Hybrid Low-Rank Natural Gradient Descent MethodBaorun Mu, Saeed Soori, Bugra Can, Mert Gürbüzbalaban 等SC 2022 · 被引用 3 次
- A Layer-Wise Natural Gradient Optimizer for Training Deep Neural NetworksXiaolei Liu, Shaoshuai Li, Kaixin Gao, Binfeng WangNeurIPS 2024 · 被引用 2 次
