Optimizing Distributed ML Communication with Fused Computation-Collective Operations
Kishore Punniyamurthy, Khaled Hamidouche, Bradford M. Beckmann
摘要
In order to satisfy their ever increasing capacity and compute requirements, machine learning models are distributed across multiple nodes using numerous parallelism strategies. As a result, collective communications are often on the critical path, and hiding their latency by overlapping kernel-granular communication and computation is difficult due to the absence of independent computation.
In this work, we propose fusing computation with dependent collective communication by leveraging GPUs' massive parallelism and GPU-initiated communication. We have developed selfcontained GPU kernels where threadblocks/workgroups (WGs) immediately communicate their results to remote GPUs when they complete their computation. Meanwhile, other WGs within the same kernel perform overlapping computation, maintaining high ALU utilization. Such fine-grain overlapping provides the additional benefit that peak network bandwidth demand is reduced and communication is spread across the entire lifetime of application rather than only at kernel boundaries. Furthermore, we propose zero-copy optimizations for scale-up communication where the data computed by one GPU is directly communicated to peer GPUs, eliminating intermediate stores and buffering.
We demonstrate our approach by creating three prototype fused operators (embedding + All-to-All, GEMV + AllReduce, and GEMM + All-to-All) to address the pervasive communication overheads observed in deep learning recommendation models (DLRM), Transformers and Mixture of Experts (MoE) model architectures. In order to demonstrate that our approach can be integrated into ML frameworks for wide adoption in production environments, we expose our fused operators as new PyTorch operators as well as extend the Triton framework to enable them. Our evaluations show that our approach can effectively overlap communication with computations, subsequently reducing their combined execution time than the current collective library-based approaches. Our scale-up GEMV + AllReduce and GEMM + Allto-All implementations achieve 13% (up to 22%) and 12% (up to 20%) lower execution time, while our fused embedding + All-to-All reduces execution time by 20% and 31% for intra-node and inter-node configurations. Large scale-out simulations indicate that our approach reduces DLRM execution time by 21% for 128 node system.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- FlashMoE: Fast Distributed MoE in a Single KernelOsayamen Jonathan Aimuyo, Byungsoo Oh, Rachee SinghNeurIPS 2025 · 被引用 22 次
- T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & CollectivesSuchita Pati, Shaizeen Aga, Mahzabeen Islam, Nuwan Jayasena 等ASPLOS 2024 · 被引用 13 次
- FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts ModelsXinglin Pan, Wenxiang Lin, Lin Zhang, Shaohuai Shi 等ASPLOS 2025 · 被引用 12 次
- Chimera: Communication Fusion for Hybrid Parallelism in Large Language ModelsLe Qin, Junwei Cui, Weilin Cai, Jiayi HuangISCA 2025 · 被引用 11 次
- FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts TrainingYunqi Gao, Bing Hu, Mahdi Boloursaz Mashhadi, A-Long Jin 等NeurIPS 2025 · 被引用 6 次
它引用的顶会 Paper10
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleSamyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang 等ICML 2022 · 被引用 523 次
- Splitwise: Efficient Generative LLM Inference Using Phase SplittingPratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah 等ISCA 2024 · 被引用 282 次
- Breaking the computation and communication abstraction barrier in distributed machine learning workloadsAbhinav Jangda, Jun Huang, Guodong Liu, Amir Hossein Nodehi Sabet 等ASPLOS 2022 · 被引用 68 次
- Synthesizing optimal collective algorithmsZixian Cai, Zhengyang Liu, Saeed Maleki, Madanlal Musuvathi 等PPoPP 2021 · 被引用 64 次
- Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning ModelsShibo Wang, Jinliang Wei, Amit Sabne, Andy Davis 等ASPLOS 2023 · 被引用 64 次
相关 Paper
- Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication OverlapXinwei Qiang, Yue Guan, Zhengding Hu, Keren Zhou 等OSDI 2026 · 被引用 3 次
- RoCC: Harnessing Raster Operations Pipeline for Efficient Tensor Collective CommunicationYuan Feng, Daniel Wong, Hyeran JeonISCA 2026
- ARK: GPU-driven Code Execution for Distributed Deep LearningChangho Hwang, KyoungSoo Park, Ran Shu, Xinyuan Qu 等NSDI 2023 · 被引用 22 次
- Libra: Contention-Aware GPU Thread Allocation for Data Parallel Training in High Speed NetworksYunzhuo Liu, Bo Jiang, Shizhen Zhao, Tao Lin 等INFOCOM 2023 · 被引用 4 次
- Preemptive All-reduce Scheduling for Expediting Distributed DNN TrainingYixin Bao, Yanghua Peng, Yangrui Chen, Chuan WuINFOCOM 2020 · 被引用 67 次
