Accelerating Distributed K-FAC with Efficient Collective Communication and Scheduling
Lin Zhang, Shaohuai Shi, Bo Li
Abstract
Kronecker-factored approximate curvature (K-FAC) has been shown to achieve faster convergence than SGD in training deep neural networks. However, existing distributed K-FAC (D-KFAC) relies on the all-reduce collective for communications and scheduling, which incurs excessive communications in each iteration. In this work, we propose a new form of D-KFAC with a reduce-based alternative to eliminate redundant communications. This poses new challenges and opportunities in that the reduce collective requires a root worker to collect the results, which considerably complicates the communication scheduling. To this end, we formulate an optimization problem that determines tensor fusion and tensor placement simultaneously aiming to minimize the training iteration time. We develop novel communication scheduling strategies and propose a placement-aware D-KFAC (PAD-KFAC) training algorithm, which further improves communication efficiency. Our experimental results on a 64-GPU cluster interconnected with 10Gb/s and 100Gb/s Ethernet show that our PAD-KFAC can achieve an average of 27% and 17% improvement over state-of-the-art D-KFAC methods, respectively.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 96a1f1ba-f709-424d-9bbf-d4a4a486c392Cited by top-tier papers1
Ask how each one uses itRelated papers
- Exploiting Simultaneous Communications to Accelerate Data Parallel Distributed Deep LearningShaohuai Shi, Xiaowen Chu, Bo LiINFOCOM 2021 · 36 citations
- Convolutional neural network training with distributed K-FACJ. Gregory Pauloski, Zhao Zhang, Lei Huang, Weijia Xu et al.SC 2020 · 26 citations
- Preemptive All-reduce Scheduling for Expediting Distributed DNN TrainingYixin Bao, Yanghua Peng, Yangrui Chen, Chuan WuINFOCOM 2020 · 67 citations
- SKFAC: Training Neural Networks With Faster Kronecker-Factored Approximate CurvatureZedong Tang, Fenlong Jiang, Maoguo Gong, Hao Li et al.CVPR 2021
- Libra: Contention-Aware GPU Thread Allocation for Data Parallel Training in High Speed NetworksYunzhuo Liu, Bo Jiang, Shizhen Zhao, Tao Lin et al.INFOCOM 2023 · 4 citations
