Accelerating Collective Communication in Data Parallel Training across Deep Learning Frameworks
Joshua Romero, Junqi Yin, Nouamane Laanait, Bing Xie, M. Todd Young, Sean Treichler, Vitalii Starchenko, Albina Y. Borisevich, Alex Sergeev, Michael A. Matheson
Abstract
This work develops new techniques within Horovod, a generic communication library supporting data parallel training across deep learning frameworks. In particular, we improve the Horovod control plane by implementing a new coordination scheme that takes advantage of the characteristics of the typical data parallel training paradigm, namely the repeated execution of collectives on the gradients of a fixed set of tensors. Using a caching strategy, we execute Horovod’s existing coordinator-worker logic only once during a typical training run, replacing it with a more efficient decentralized orchestration strategy using the cached data and a global intersection of a bitvector for the remaining training duration. Next, we introduce a feature for end users to explicitly group collective operations, enabling finer grained control over the communication buffer sizes. To evaluate our proposed strategies, we conduct experiments on a world-class supercomputer — Summit. We compare our proposals to Horovod’s original design and observe 2x performance improvement at a scale of 6000 GPUs; we also compare them against tf.distribute and torch.DDP and achieve 12% better and comparable performance, respectively, using up to 1536 GPUs; we compare our solution against BytePS in typical HPC settings and achieve about 20% better performance on a scale of 768 GPUs. Finally, we test our strategies on a scientific application (STEMDL) using up to 27,600 GPUs (the entire Summit) and show that we achieve a near-linear scaling of 0.93 with a sustained performance of 1.54 exaflops (with standard error +- 0.02) in FP16 precision.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a38690f5-95bb-4445-8b13-ad1ba9ce3d2eCited by top-tier papers6
- High-throughput and Flexible Host Networking for Accelerated ComputingAthinagoras Skiadopoulos, Zhiqiang Xie, Mark Zhao, Qizhe Cai et al.OSDI 2024 · 11 citations
- ResCCL: Resource-Efficient Scheduling for Collective CommunicationTongrui Liu, Chenyang Hei, Fuliang Li, Chengxi Gao et al.SIGCOMM 2025 · 11 citations
- PIMnet: A Domain-Specific Network for Efficient Collective Communication in Scalable PIMHyojun Son, Gilbert Jonatan, Xiangyu Wu, Haeyoon Cho et al.HPCA 2025 · 7 citations
- DES-LOC: Desynced Low Communication Adaptive Optimizers for Foundation ModelsAlex Iacob, Lorenzo Sani, Mher Safaryan, Paris Giampouras et al.ICLR 2026 · 2 citations
- FusedRec: Fused Embedding Communication for Distributed Recommendation Training on GPUsXuanteng Huang, Fan Li, Riyang Hu, Jianchang Zhang et al.AAAI 2026 · 1 citation
Builds on2
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi et al.OSDI 2020 · 390 citations
- Decentralized Deep Learning with Arbitrary Communication CompressionAnastasia Koloskova, Tao Lin, Sebastian U. Stich, Martin JaggiICLR 2020 · 263 citations
Related papers
- BAGUA: Scaling up Distributed Learning with System RelaxationsShaoduo Gan, Xiangru Lian, Rui Wang, Jianbin Chang et al.VLDB 2022 · 35 citations
- Libra: Contention-Aware GPU Thread Allocation for Data Parallel Training in High Speed NetworksYunzhuo Liu, Bo Jiang, Shizhen Zhao, Tao Lin et al.INFOCOM 2023 · 4 citations
- Exploiting Simultaneous Communications to Accelerate Data Parallel Distributed Deep LearningShaohuai Shi, Xiaowen Chu, Bo LiINFOCOM 2021 · 36 citations
- Preemptive All-reduce Scheduling for Expediting Distributed DNN TrainingYixin Bao, Yanghua Peng, Yangrui Chen, Chuan WuINFOCOM 2020 · 67 citations
- Enabling Efficient Large-Scale Deep Learning Training with Cache Coherent Disaggregated Memory SystemsZixuan Wang, Joonseop Sim, Euicheol Lim, Jishen ZhaoHPCA 2022 · 9 citations
