Inefficiency of K-FAC for Large Batch Size Training
Linjian Ma, Gabe Montague, Jiayu Ye, Zhewei Yao, Amir Gholami, Kurt Keutzer, Michael W. Mahoney
摘要
In stochastic optimization, using large batch sizes during training can leverage parallel resources to produce faster wall-clock training times per training epoch. However, for both training loss and testing error, recent results analyzing large batch Stochastic Gradient Descent (SGD) have found sharp diminishing returns, beyond a certain critical batch size. In the hopes of addressing this, it has been suggested that the Kronecker-Factored Approximate Curvature (K-FAC) method allows for greater scalability to large batch sizes, for non-convex machine learning problems such as neural network optimization, as well as greater robustness to variation in model hyperparameters. Here, we perform a detailed empirical analysis of large batch size training for both K-FAC and SGD, evaluating performance in terms of both wall-clock time and aggregate computational cost. Our main results are twofold: first, we find that both K-FAC and SGD doesn't have ideal scalability behavior beyond a certain batch size, and that K-FAC does not exhibit improved large-batch scalability behavior, as compared to SGD; and second, we find that K-FAC, in addition to requiring more hyperparameters to tune, suffers from similar hyperparameter sensitivity behavior as does SGD. We discuss extensive results using ResNet and AlexNet on CIFAR-10 and SVHN, respectively, as well as more general implications of our findings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- LocFedMix-SL: Localize, Federate, and Mix for Improved Scalability, Convergence, and Latency in Split LearningSeungeun Oh, Jihong Park, Praneeth Vepakomma, Sihun Baek 等WWW 2022 · 被引用 66 次
- Distributed Learning of Fully Connected Neural Networks using Independent Subnet TrainingBinhang Yuan, Cameron R. Wolfe, Chen Dun, Yuxin Tang 等VLDB 2022 · 被引用 42 次
- Convolutional neural network training with distributed K-FACJ. Gregory Pauloski, Zhao Zhang, Lei Huang, Weijia Xu 等SC 2020 · 被引用 26 次
- Second-Order Neural ODE OptimizerGuan-Horng Liu, Tianrong Chen, Evangelos A. TheodorouNeurIPS 2021 · 被引用 20 次
- HyLo: A Hybrid Low-Rank Natural Gradient Descent MethodBaorun Mu, Saeed Soori, Bugra Can, Mert Gürbüzbalaban 等SC 2022 · 被引用 3 次
相关 Paper
- KAISA: an adaptive second-order optimizer framework for deep neural networksJ. Gregory Pauloski, Qi Huang, Lei Huang, Shivaram Venkataraman 等SC 2021 · 被引用 14 次
- SKFAC: Training Neural Networks With Faster Kronecker-Factored Approximate CurvatureZedong Tang, Fenlong Jiang, Maoguo Gong, Hao Li 等CVPR 2021
- Gradient Descent on Neurons and its Link to Approximate Second-order OptimizationFrederik BenzingICML 2022 · 被引用 31 次
- Kronecker-Factored Approximate Curvature for Modern Neural Network ArchitecturesRuna Eschenhagen, Alexander Immer, Richard E. Turner, Frank Schneider 等NeurIPS 2023 · 被引用 62 次
- Accelerating Distributed K-FAC with Efficient Collective Communication and SchedulingLin Zhang, Shaohuai Shi, Bo LiINFOCOM 2023 · 被引用 4 次
