Multipath Collective Communication Beyond Scale-up Networks in GPU Clouds
Yuchen Xu, Jianglong Nie, Baojia Li, Mingzhuo Chen, Hao Lu, Guanyu Qu, Zhenchuan Liu, Shuangshuang Yin, Xiaojie Huang, Chunzhi He, Yinben Xia, Quan Wen
Abstract
Hardware vendors introduce scale-up networks interconnecting accelerators to speed up communication in distributed training. We argue that in the context of GPU clouds, the scale-out network can be a good complement to the collective communication that is originally performed on scale-up networks. We build a system named MPCCS to enable multipath transmission of collectives on both networks. MPCCS essentially splits the traffic of collective flows to two networks. MPCCS overcomes three challenges caused by the progress of hardware-offloaded networks and the diversity of collective communication. First, it devises a dual-window protocol to enable the runtime traffic splitting over rigid hardware-offloaded interfaces. Second, it devises a bandwidth-delay product (BDP) estimation algorithm to enable bandwidth-adaptive traffic splitting, overcoming the difficulty of invisibility of transmission states (RTT and throughput) due to hardware encapsulation. Third, it devises coflow-synchronized multipath transmission for collective flows, which achieves universal applicability for diverse collectives in terms of correctness and performance optimality. We implement MPCCS and conduct extensive experiments both on the testbed and in simulation. MPCCS achieves a 23% 54% speedup compared with vanilla NCCL on communication microbenchmarks and above 10% acceleration for LLM training.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- MCCS: A Service-based Approach to Collective Communication for Multi-Tenant CloudYongji Wu, Yechen Xu, Jingrong Chen, Zhaodong Wang et al.SIGCOMM 2024 · 15 citations
- RoCC: Harnessing Raster Operations Pipeline for Efficient Tensor Collective CommunicationYuan Feng, Daniel Wong, Hyeran JeonISCA 2026
- COCCL: A Collective Communication Library Supporting Easy Integration and Configuration of Customized Compression for Scalable LLM TrainingXingchen Liu, Haoran Kong, Hairui Zhao, Shengkai Lyu et al.PPoPP 2026 · 3 citations
- An In-Network Architecture for Accelerating Shared-Memory Multiprocessor CollectivesBenjamin Klenk, Nan Jiang, Greg Thorson, Larry DennisonISCA 2020 · 67 citations
- Disdp: Disaggregating Compute, Network, and Storage for Model-Sharded Data-Parallel TrainingMo Sun, Zihan Yang, Changyue Liao, Yingtao Li et al.ISCA 2026 · 1 citation
