SC2022Top-tier venue
HyLo: A Hybrid Low-Rank Natural Gradient Descent Method
Baorun Mu, Saeed Soori, Bugra Can, Mert Gürbüzbalaban, Maryam Mehri Dehnavi
Abstract
This work presents a Hybrid Low-Rank Natural Gradient Descent method, called HyLo, that accelerates the training time of deep neural networks. Natural gradient descent (NGD) requires computing the inverse of the Fisher information matrix (FIM), which is typically expensive at largescale. Kronecker factorization methods such as K F A C attempt to improve NGD's running time by approximating the FIM with Kronecker factors. However, the size of Kronecker factors increases quadratically as the model size grows. Instead, in HyLo, we use the Sherman-Morrison-Woodbury variant of NGD (SNGD) and propose a reformulation of SNGD to resolve its scalability issues. HyL o uses a computationally-efficient low-rank factorization to achieve superior timing for Fisher inverses. We evaluate HyL o on large models including ResNet-50, U-Net, and ResNet-32 on up to 64 GPUs. H yL o converges 1.4×-2.1× faster than the state-of-the-art distributed implementation of K FA C and reduces the computation and communication time up to 350× and 10.7× on ResNet-50.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9497c9db-5794-49f8-9b87-8c2cd81670b9Cited by top-tier papers1
Ask how each one uses itBuilds on9
- MLPerf Inference BenchmarkVijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson et al.ISCA 2020 · 517 citations
- WoodFisher: Efficient Second-Order Approximation for Neural Network CompressionSidak Pal Singh, Dan AlistarhNeurIPS 2020 · 217 citations
- Practical Quasi-Newton Methods for Training Deep Neural NetworksDonald Goldfarb, Yi Ren, Achraf BahamouNeurIPS 2020 · 130 citations
- BackPACK: Packing more into BackpropFelix Dangel, Frederik Kunstner, Philipp HennigICLR 2020 · 114 citations
- Understanding Approximate Fisher Information for Fast Convergence of Natural Gradient Descent in Wide Neural NetworksRyo Karakida, Kazuki OsawaNeurIPS 2020 · 39 citations
Related papers
- SKFAC: Training Neural Networks With Faster Kronecker-Factored Approximate CurvatureZedong Tang, Fenlong Jiang, Maoguo Gong, Hao Li et al.CVPR 2021
- THOR, Trace-based Hardware-driven Layer-Oriented Natural Gradient Descent ComputationMengyun Chen, Kai-Xin Gao, Xiaolei Liu, Zidong Wang et al.AAAI 2021 · 7 citations
- Convolutional neural network training with distributed K-FACJ. Gregory Pauloski, Zhao Zhang, Lei Huang, Weijia Xu et al.SC 2020 · 26 citations
- A Layer-Wise Natural Gradient Optimizer for Training Deep Neural NetworksXiaolei Liu, Shaoshuai Li, Kaixin Gao, Binfeng WangNeurIPS 2024 · 2 citations
- Rich Information is Affordable: A Systematic Performance Analysis of Second-order Optimization Using K-FACYuichiro Ueno, Kazuki Osawa, Yohei Tsuji, Akira Naruse et al.KDD 2020 · 9 citations
