SC2020Top-tier venue
Convolutional neural network training with distributed K-FAC
J. Gregory Pauloski, Zhao Zhang, Lei Huang, Weijia Xu, Ian T. Foster
Abstract
Training neural networks with many processors can reduce time-to-solution; however, it is challenging to maintain convergence and efficiency at large scales. The Kroneckerfactored Approximate Curvature (K-FAC) was recently proposed as an approximation of the Fisher Information Matrix that can be used in natural gradient optimizers. We investigate here a scalable K-FAC design and its applicability in convolutional neural network (CNN) training at scale. We study optimization techniques such as layer-wise distribution strategies, inverse-free second-order gradient evaluation, and dynamic K-FAC update decoupling to reduce training time while preserving convergence. We use residual neural networks (ResNet) applied to the CIFAR-10 and ImageNet-1k datasets to evaluate the correctness and scalability of our K-FAC gradient preconditioner. With ResNet-50 on the ImageNet-1k dataset, our distributed K-FAC implementation converges to the 75.9% MLPerf baseline in 18-25% less time than does the classic stochastic gradient descent (SGD) optimizer across scales on a GPU cluster.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1668c106-7ef8-4868-8e79-3b3cb0495c74Cited by top-tier papers9
- COMET: A Novel Memory-Efficient Deep Learning Training Framework by Using Error-Bounded Lossy CompressionSian Jin, Chengming Zhang, Xintong Jiang, Yunhe Feng et al.VLDB 2022 · 39 citations
- An Improved Empirical Fisher Approximation for Natural Gradient DescentXiaodong Wu, Wenyi Yu, Chao Zhang, Philip C. WoodlandNeurIPS 2024 · 27 citations
- KAISA: an adaptive second-order optimizer framework for deep neural networksJ. Gregory Pauloski, Qi Huang, Lei Huang, Shivaram Venkataraman et al.SC 2021 · 14 citations
- A Trace-restricted Kronecker-Factored Approximation to Natural GradientKai-Xin Gao, Xiao-Lei Liu, Zheng-Hai Huang, Min Wang et al.AAAI 2021 · 13 citations
- An Oracle for Guiding Large-Scale Model/Hybrid Parallel Training of Convolutional Neural NetworksAlbert Njoroge Kahira, Truong Thao Nguyen, Leonardo Bautista-Gomez, Ryousei Takano et al.HPDC 2021 · 11 citations
Builds on1
Related papers
- SKFAC: Training Neural Networks With Faster Kronecker-Factored Approximate CurvatureZedong Tang, Fenlong Jiang, Maoguo Gong, Hao Li et al.CVPR 2021
- Rich Information is Affordable: A Systematic Performance Analysis of Second-order Optimization Using K-FACYuichiro Ueno, Kazuki Osawa, Yohei Tsuji, Akira Naruse et al.KDD 2020 · 9 citations
- HyLo: A Hybrid Low-Rank Natural Gradient Descent MethodBaorun Mu, Saeed Soori, Bugra Can, Mert Gürbüzbalaban et al.SC 2022 · 3 citations
- Gradient Descent on Neurons and its Link to Approximate Second-order OptimizationFrederik BenzingICML 2022 · 31 citations
- Accelerating Distributed K-FAC with Efficient Collective Communication and SchedulingLin Zhang, Shaohuai Shi, Bo LiINFOCOM 2023 · 4 citations
