MSCCL++: Rethinking GPU Communication Abstractions for AI Inference
Changho Hwang, Peng Cheng, Roshan Dathathri, Abhinav Jangda, Saeed Maleki, Madan Musuvathi, Olli Saarikivi, Aashaka Shah, Ziyue Yang, Binyang Li, Caio Rocha, Qinghua Zhou
Abstract
AI applications increasingly run on fast-evolving, heterogeneous hardware to maximize performance, but general-purpose libraries lag in supporting these features. Performance-minded programmers often build custom communication stacks that are fast but error-prone and non-portable. This paper introduces MSCCL++, a design methodology for developing high-performance, portable communication kernels. It provides (1) a low-level, performance-preserving primitive interface that exposes minimal hardware abstractions while hiding the complexities of synchronization and consistency, (2) a higher-level DSL for application developers to implement workload-specific communication algorithms, and (3) a library of efficient algorithms implementing the standard collective API, enabling adoption by users with minimal expertise. Compared to state-of-the-art baselines, MSCCL++ achieves geomean speedups of 1.7× (up to 5.4×) for collective communication and 1.2× (up to 1.38×) for AI inference workloads. MSCCL++ is in production of multiple AI services provided by Microsoft Azure, and has also been adopted by RCCL, the GPU collective communication library maintained by AMD. MSCCL++ is open source and available at https://github.com/microsoft/mscclpp. Our two years of experience with MSCCL++ suggests that its abstractions are robust, enabling support for new hardware features, such as multimem, within weeks of development.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Kareus: Joint Reduction of Dynamic and Static Energy in Large Model TrainingRuofan Wu, Jae-Won Chung, Mosharaf ChowdhuryOSDI 2026 · 8 citations
- UBEP: Re-architecting Expert Parallelism Communication Library for Production SuperpodsYipeng Liu, Chang Liu, Si Shen, Jiaqi Zheng et al.SIGCOMM 2026
Builds on10
- Splitwise: Efficient Generative LLM Inference Using Phase SplittingPratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah et al.ISCA 2024 · 282 citations
- Breaking the computation and communication abstraction barrier in distributed machine learning workloadsAbhinav Jangda, Jun Huang, Guodong Liu, Amir Hossein Nodehi Sabet et al.ASPLOS 2022 · 68 citations
- Synthesizing optimal collective algorithmsZixian Cai, Zhengyang Liu, Saeed Maleki, Madanlal Musuvathi et al.PPoPP 2021 · 64 citations
- Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning ModelsShibo Wang, Jinliang Wei, Amit Sabne, Andy Davis et al.ASPLOS 2023 · 64 citations
- Crux: GPU-Efficient Communication Scheduling for Deep Learning TrainingJiamin Cao, Yu Guan, Kun Qian, Jiaqi Gao et al.SIGCOMM 2024 · 60 citations
Related papers
- MSCCLang: Microsoft Collective Communication LanguageMeghan Cowan, Saeed Maleki, Madanlal Musuvathi, Olli Saarikivi et al.ASPLOS 2023 · 40 citations
- TCCL: Discovering Better Communication Paths for PCIe GPU ClustersHeehoon Kim, Junyeol Ryu, Jaejin LeeASPLOS 2024 · 26 citations
- MCCS: A Service-based Approach to Collective Communication for Multi-Tenant CloudYongji Wu, Yechen Xu, Jingrong Chen, Zhaodong Wang et al.SIGCOMM 2024 · 15 citations
- SyCCL: Exploiting Symmetry for Efficient Collective Communication SchedulingJiamin Cao, Shangfeng Shi, Jiaqi Gao, Weisen Liu et al.SIGCOMM 2025 · 15 citations
- ACCL+: an FPGA-Based Collective Engine for Distributed ApplicationsZhenhao He, Dario Korolija, Yu Zhu, Benjamin Ramhorst et al.OSDI 2024 · 12 citations
