Multipath Collective Communication Beyond Scale-up Networks in GPU Clouds
Yuchen Xu, Jianglong Nie, Baojia Li, Mingzhuo Chen, Hao Lu, Guanyu Qu, Zhenchuan Liu, Shuangshuang Yin, Xiaojie Huang, Chunzhi He, Yinben Xia, Quan Wen
摘要
Hardware vendors introduce scale-up networks interconnecting accelerators to speed up communication in distributed training. We argue that in the context of GPU clouds, the scale-out network can be a good complement to the collective communication that is originally performed on scale-up networks. We build a system named MPCCS to enable multipath transmission of collectives on both networks. MPCCS essentially splits the traffic of collective flows to two networks. MPCCS overcomes three challenges caused by the progress of hardware-offloaded networks and the diversity of collective communication. First, it devises a dual-window protocol to enable the runtime traffic splitting over rigid hardware-offloaded interfaces. Second, it devises a bandwidth-delay product (BDP) estimation algorithm to enable bandwidth-adaptive traffic splitting, overcoming the difficulty of invisibility of transmission states (RTT and throughput) due to hardware encapsulation. Third, it devises coflow-synchronized multipath transmission for collective flows, which achieves universal applicability for diverse collectives in terms of correctness and performance optimality. We implement MPCCS and conduct extensive experiments both on the testbed and in simulation. MPCCS achieves a 23% 54% speedup compared with vanilla NCCL on communication microbenchmarks and above 10% acceleration for LLM training.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- MCCS: A Service-based Approach to Collective Communication for Multi-Tenant CloudYongji Wu, Yechen Xu, Jingrong Chen, Zhaodong Wang 等SIGCOMM 2024 · 被引用 15 次
- RoCC: Harnessing Raster Operations Pipeline for Efficient Tensor Collective CommunicationYuan Feng, Daniel Wong, Hyeran JeonISCA 2026
- COCCL: A Collective Communication Library Supporting Easy Integration and Configuration of Customized Compression for Scalable LLM TrainingXingchen Liu, Haoran Kong, Hairui Zhao, Shengkai Lyu 等PPoPP 2026 · 被引用 3 次
- An In-Network Architecture for Accelerating Shared-Memory Multiprocessor CollectivesBenjamin Klenk, Nan Jiang, Greg Thorson, Larry DennisonISCA 2020 · 被引用 67 次
- Disdp: Disaggregating Compute, Network, and Storage for Model-Sharded Data-Parallel TrainingMo Sun, Zihan Yang, Changyue Liao, Yingtao Li 等ISCA 2026 · 被引用 1 次
