MSCCLang: Microsoft Collective Communication Language
Meghan Cowan, Saeed Maleki, Madanlal Musuvathi, Olli Saarikivi, Yifan Xiong
摘要
Machine learning models with millions or billions of parameters are increasingly trained and served on large multi-GPU systems. As models grow in size and execute on more GPUs, collective communication becomes a bottleneck. Custom collective algorithms optimized for both particular network topologies and applicationspecific communication patterns can alleviate this bottleneck and help these applications scale. However, implementing correct and efficient custom algorithms is challenging.
This paper introduces MSCCLang, a system for programmable GPU communication. MSCCLang provides a domain specific language for writing collective communication algorithms and an optimizing compiler for lowering them to an executable form, which can be executed efficiently and flexibly in an interpreter-based runtime. We used MSCCLang to write novel collective algorithms for AllReduce and AllToAll that are up to 1.9× and 1.3× faster than hand-optimized implementations, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- TACOS: Topology-Aware Collective Algorithm Synthesizer for Distributed Machine LearningWilliam Won, Midhilesh Elavazhagan, Sudarshan Srinivasan, Swati Gupta 等MICRO 2024 · 被引用 36 次
- AutoCCL: Automated Collective Communication Tuning for Accelerating Distributed and Parallel DNN TrainingGuanbin Xu, Zhihao Le, Yinhe Chen, Zhiqi Lin 等NSDI 2025 · 被引用 27 次
- GREYHOUND: Hunting Fail-Slows in Hybrid-Parallel Training at ScaleTianyuan Wu, Wei Wang, Yinghao Yu, Siran Yang 等USENIX ATC 2025 · 被引用 19 次
- Octopus: Enhancing CXL Memory Pods via Sparse TopologyYuhong Zhong, Fiodar Kazhamiaka, Pantea Zardoshti, Shuwei Teng 等NSDI 2026 · 被引用 15 次
- PID-Comm: A Fast and Flexible Collective Communication Framework for Commodity Processing-in-DIMM DevicesSi Ung Noh, Junguk Hong, Chaemin Lim, Seongyeon Park 等ISCA 2024 · 被引用 12 次
它引用的顶会 Paper4
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi 等OSDI 2020 · 被引用 390 次
- ATP: In-network Aggregation for Multi-tenant LearningChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen 等NSDI 2021 · 被引用 359 次
- Synthesizing optimal collective algorithmsZixian Cai, Zhengyang Liu, Saeed Maleki, Madanlal Musuvathi 等PPoPP 2021 · 被引用 64 次
- Scaling Distributed Machine Learning with In-Network AggregationAmedeo Sapio, Marco Canini, Chen-Yu Ho, Jacob Nelson 等NSDI 2021
相关 Paper
- TACCL: Guiding Collective Algorithm Synthesis using Communication SketchesAashaka Shah, Vijay Chidambaram, Meghan Cowan, Saeed Maleki 等NSDI 2023
- TCCL: Discovering Better Communication Paths for PCIe GPU ClustersHeehoon Kim, Junyeol Ryu, Jaejin LeeASPLOS 2024 · 被引用 26 次
- OptCCL: Scalable Synthesis of Optimal Collective Communication AlgorithmsRichard Shapley, Rachit Agarwal, David B. ShmoysSIGCOMM 2026
- MSCCL++: Rethinking GPU Communication Abstractions for AI InferenceChangho Hwang, Peng Cheng, Roshan Dathathri, Abhinav Jangda 等ASPLOS 2026 · 被引用 5 次
- ResCCL: Resource-Efficient Scheduling for Collective CommunicationTongrui Liu, Chenyang Hei, Fuliang Li, Chengxi Gao 等SIGCOMM 2025 · 被引用 11 次
