Efficient all-to-all Collective Communication Schedules for Direct-connect Topologies
Prithwish Basu, Liangyu Zhao, Jason Fantl, Siddharth Pal, Arvind Krishnamurthy, Joud Khoury
摘要
The all-to-all collective communications primitive is widely used in machine learning (ML) and high performance computing (HPC) workloads, and optimizing its performance is of interest to both ML and HPC communities. All-to-all is a particularly challenging workload that can severely strain the underlying interconnect bandwidth at scale. This paper takes a holistic approach to optimize the performance of all-to-all collective communications on supercomputer-scale direct-connect interconnects. We address several algorithmic and practical challenges in developing efficient and bandwidth-optimal all-to-all schedules for any topology and lowering the schedules to various runtimes and interconnect technologies. We also propose a novel topology that delivers near-optimal all-to-all performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- MCCS: A Service-based Approach to Collective Communication for Multi-Tenant CloudYongji Wu, Yechen Xu, Jingrong Chen, Zhaodong Wang 等SIGCOMM 2024 · 被引用 15 次
- Efficient Direct-Connect Topologies for Collective CommunicationsLiangyu Zhao, Siddharth Pal, Tapan Chugh, Weiyang Wang 等NSDI 2025
它引用的顶会 Paper10
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- Jupiter evolving: transforming google's datacenter network via optical circuit switches and software-defined networkingLeon Poutievski, Omid Mashayekhi, Joon Ong, Arjun Singh 等SIGCOMM 2022 · 被引用 230 次
- TopoOpt: Co-optimizing Network Topology and Parallelization Strategy for Distributed Training JobsWeiyang Wang, Moein Khazraee, Zhizhen Zhong, Manya Ghobadi 等NSDI 2023 · 被引用 215 次
- Expanding across time to deliver bandwidth efficiency and low latencyWilliam M. Mellette, Rajdeep Das, Yibo Guo, Rob McGuinness 等NSDI 2020 · 被引用 194 次
- SiP-ML: high-bandwidth optical network interconnects for machine learning trainingMehrdad Khani Shirkoohi, Manya Ghobadi, Mohammad Alizadeh, Ziyi Zhu 等SIGCOMM 2021 · 被引用 94 次
相关 Paper
- SyCCL: Exploiting Symmetry for Efficient Collective Communication SchedulingJiamin Cao, Shangfeng Shi, Jiaqi Gao, Weisen Liu 等SIGCOMM 2025 · 被引用 15 次
- FAST: An Efficient Scheduler for All-to-All GPU CommunicationYiran Lei, Dongjoo Lee, Liangyu Zhao, Daniar Kurniawan 等NSDI 2026 · 被引用 13 次
- ForestColl: Throughput-Optimal Collective Communications on Heterogeneous Network FabricsLiangyu Zhao, Saeed Maleki, Yuanhong Wang, Zezhou Wang 等NSDI 2026
- OptCCL: Scalable Synthesis of Optimal Collective Communication AlgorithmsRichard Shapley, Rachit Agarwal, David B. ShmoysSIGCOMM 2026
- Harvest: Adaptive Photonic Switching Schedules for Collective Communication in Scale-up DomainsMahir Rahman, Samuel Joseph, Nihar Kodkani, Behnaz Arzani 等SIGCOMM 2026
