Preemptive All-reduce Scheduling for Expediting Distributed DNN Training
Yixin Bao, Yanghua Peng, Yangrui Chen, Chuan Wu
Abstract
Data-parallel training is widely used for scaling DNN training over large datasets, using the parameter server or all-reduce architecture. Communication scheduling has been promising to accelerate distributed DNN training, which aims to overlap communication with computation by scheduling the order of communication operations. We identify two limitations of previous communication scheduling work. First, layer-wise computation graph has been a common assumption, while modern machine learning frameworks (e.g., TensorFlow) use a sophisticated directed acyclic graph (DAG) representation as the execution model. Second, the default sizes of tensors are often less than optimal for transmission scheduling and bandwidth utilization. We propose PACE, a communication scheduler that preemptively schedules (potentially fused) all-reduce tensors based on the DAG of DNN training, guaranteeing maximal overlapping of communication with computation and high bandwidth utilization. The scheduler contains two integrated modules: given a DAG, we identify the best tensor-preemptive communication schedule that minimizes the training time; exploiting the optimal communication scheduling as an oracle, a dynamic programming approach is developed for generating a good DAG, which merges small communication tensors for efficient bandwidth utilization. Experiments in a GPU testbed show that PACE accelerates training with representative system configurations, achieving up to 36% speed-up compared with state-of-the-art solutions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c66b8954-43d9-439e-a130-5b8cda242ca8Cited by top-tier papers9
- MegaScale: Scaling Large Language Model Training to More Than 10, 000 GPUsZiheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang et al.NSDI 2024 · 415 citations
- Accelerating Distributed MoE Training and Inference with LinaJiamin Li, Yimin Jiang, Yibo Zhu, Cong Wang et al.USENIX ATC 2023 · 191 citations
- GADGET: Online Resource Optimization for Scheduling Ring-All-Reduce Learning JobsMenglu Yu, Ye Tian, Bo Ji, Chuan Wu et al.INFOCOM 2022 · 38 citations
- SHADE: Enable Fundamental Cacheability for Distributed Deep Learning TrainingRedwan Ibne Seraj Khan, Ahmad Hossein Yazdani, Yuqi Fu, Arnab K. Paul et al.FAST 2023 · 29 citations
- SAPipe: Staleness-Aware Pipeline for Data Parallel DNN TrainingYangrui Chen, Cong Xie, Meng Ma, Juncheng Gu et al.NeurIPS 2022 · 24 citations
Related papers
- Near-Optimal Topology-adaptive Parameter Synchronization in Distributed DNN TrainingZhe Zhang, Chuan Wu, Zongpeng LiINFOCOM 2021 · 14 citations
- Libra: Contention-Aware GPU Thread Allocation for Data Parallel Training in High Speed NetworksYunzhuo Liu, Bo Jiang, Shizhen Zhao, Tao Lin et al.INFOCOM 2023 · 4 citations
- Exploiting Simultaneous Communications to Accelerate Data Parallel Distributed Deep LearningShaohuai Shi, Xiaowen Chu, Bo LiINFOCOM 2021 · 36 citations
- Mercury: A Simple Transport Layer Scheduler to Accelerate Distributed DNN TrainingQingyang Duan, Zeqin Wang, Yuedong Xu, Shaoteng Liu et al.INFOCOM 2022 · 12 citations
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi et al.OSDI 2020 · 390 citations
