Modeling and Optimizing the Scaling Performance in Distributed Deep Learning Training
Ting Liu, Tianhao Miao, Qinghua Wu, Zhenyu Li, Guangxin He, Jiaoren Wu, Shengzhuo Zhang, Xingwu Yang, Gareth Tyson, Gaogang Xie
摘要
Distributed Deep Learning (DDL) is widely used to accelerate deep neural network training for various Web applications. In each iteration of DDL training, each worker synchronizes neural network gradients with other workers. This introduces communication overhead and degrades the scaling performance. In this paper, we propose a recursive model, OSF (Scaling Factor considering Overlap), for estimating the scaling performance of DDL training of neural network models, given the settings of the DDL system. OSF captures two main characteristics of DDL training: the overlap between computation and communication, and the tensor fusion for batching updates. Measurements on a real-world DDL system show that OSF obtains a low estimation error (ranging from 0.5% to 8.4% for different models). Using OSF, we identify the factors that degrade the scaling performance, and propose solutions to effectively mitigate their impacts. Specifically, the proposed adaptive tensor fusion improves the scaling performance by 32.2%∼150% compared to the constant tensor fusion buffer size.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- SAFusion: Efficient Tensor Fusion with Sparsification Ahead for High-Performance Distributed DNN TrainingZhangqiang Ming, Yuchong Hu, Xinjue Zheng, Wenxiang Zhou 等HPDC 2025 · 被引用 1 次
- SSFusion: Tensor Fusion with Selective Sparsification for Efficient Distributed DNN TrainingZhangqiang Ming, Rui Wang, Yuchong Hu, Yuanhao Shu 等ICDE 2026 · 被引用 1 次
- Preemptive All-reduce Scheduling for Expediting Distributed DNN TrainingYixin Bao, Yanghua Peng, Yangrui Chen, Chuan WuINFOCOM 2020 · 被引用 67 次
- Libra: Contention-Aware GPU Thread Allocation for Data Parallel Training in High Speed NetworksYunzhuo Liu, Bo Jiang, Shizhen Zhao, Tao Lin 等INFOCOM 2023 · 被引用 4 次
- Exploiting Simultaneous Communications to Accelerate Data Parallel Distributed Deep LearningShaohuai Shi, Xiaowen Chu, Bo LiINFOCOM 2021 · 被引用 36 次
