Libra: Contention-Aware GPU Thread Allocation for Data Parallel Training in High Speed Networks
Yunzhuo Liu, Bo Jiang, Shizhen Zhao, Tao Lin, Xinbing Wang, Chenghu Zhou
Abstract
Overlapping gradient communication with backward computation is a popular technique to reduce communication cost in the widely adopted data parallel S-SGD training. However, the resource contention between computation and All-Reduce communication in GPU-based training reduces the benefits of overlap. With GPU cluster network evolving from low bandwidth TCP to high speed networks, more GPU resources are required to efficiently utilize the bandwidth, making the contention more noticeable. Existing communication libraries fail to account for such contention when allocating GPU threads and have suboptimal performance. In this paper, we propose to mitigate the contention by balancing the overlapped computation and communication time. We formulate an optimization problem that decides the communication thread allocation to reduce overall backward time. We develop a dynamic programming based near-optimal solution and extend it to co-optimize thread allocation with tensor fusion. We conduct simulated study and real-world experiment using an 8-node GPU cluster with 50Gb RDMA network training four representative DNN models. Results show that our method reduces backward time by 10%-20% compared with Horovod-NCCL, by 6%-13% compared with tensor-fusion-optimization-only methods. Simulation shows that our method achieves the best scalability with a training speedup of 1.2x over the best-performing baseline as we scale up cluster size.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Exploiting Simultaneous Communications to Accelerate Data Parallel Distributed Deep LearningShaohuai Shi, Xiaowen Chu, Bo LiINFOCOM 2021 · 36 citations
- Preemptive All-reduce Scheduling for Expediting Distributed DNN TrainingYixin Bao, Yanghua Peng, Yangrui Chen, Chuan WuINFOCOM 2020 · 67 citations
- SSFusion: Tensor Fusion with Selective Sparsification for Efficient Distributed DNN TrainingZhangqiang Ming, Rui Wang, Yuchong Hu, Yuanhao Shu et al.ICDE 2026 · 1 citation
- SAFusion: Efficient Tensor Fusion with Sparsification Ahead for High-Performance Distributed DNN TrainingZhangqiang Ming, Yuchong Hu, Xinjue Zheng, Wenxiang Zhou et al.HPDC 2025 · 1 citation
- Accelerating Distributed K-FAC with Efficient Collective Communication and SchedulingLin Zhang, Shaohuai Shi, Bo LiINFOCOM 2023 · 4 citations
