FAST: An Efficient Scheduler for All-to-All GPU Communication
Yiran Lei, Dongjoo Lee, Liangyu Zhao, Daniar Kurniawan, Chanmyeong Kim, Heetaek Jeong, Changsu Kim, Hyeonseong Choi, Liangcheng Yu, Arvind Krishnamurthy, Justine Sherry, Eriko Nurvitadhi
摘要
All-to-All(v) communication is a critical primitive in modern machine learning workloads, particularly mixture-of-experts (MoE) models. Unfortunately, efficient scheduling is challenging due to workload skew, heterogeneous two-tier fabrics, and incast congestion, compounded by the dynamic nature of MoE workloads, where traffic shifts every few hundred milliseconds. Existing schedulers are hardly scalable, incurring seconds to hours of synthesis time, making them impractical. We present FAST, an efficient All-to-All(v) scheduler. FAST addresses skew through intra-server rebalancing and enforces balanced, one-to-one scale-out transfers that avoid incast. Evaluated extensively on both NVIDIA H200 and AMD MI300X clusters, FAST consistently outperforms state-of-the-art solutions on skewed workloads while reducing synthesis time by orders of magnitude.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- TEMP: A Memory Efficient Physical-Aware Tensor Partition-Mapping Framework on Wafer-Scale ChipsHuizheng Wang, Taiquan Wei, Zichuan Wang, Dingcheng Jiang 等HPCA 2026 · 被引用 2 次
- StreamEP: Straggler-Tolerant MoE Decoding without Communication BarriersYizhuo Liang, Shaoyu Wang, Jaeyong Song, Yanqi Zhou 等SOSP 2026
- Understanding and Profiling the Accelerator Chiplet Network Using PingPointJunyeol Ryu, Ming Liu, Matthew D. SinclairSIGCOMM 2026
它引用的顶会 Paper9
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleSamyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang 等ICML 2022 · 被引用 523 次
- Accelerating Distributed MoE Training and Inference with LinaJiamin Li, Yimin Jiang, Yibo Zhu, Cong Wang 等USENIX ATC 2023 · 被引用 191 次
- RDMA over Ethernet for Distributed Training at Meta ScaleAdithya Gangidi, Rui Miao, Shengbao Zheng, Sai Jayesh Bondu 等SIGCOMM 2024 · 被引用 171 次
- Synthesizing optimal collective algorithmsZixian Cai, Zhengyang Liu, Saeed Maleki, Madanlal Musuvathi 等PPoPP 2021 · 被引用 64 次
- Rethinking Machine Learning Collective Communication as a Multi-Commodity Flow ProblemXuting Liu, Behnaz Arzani, Siva Kesava Reddy Kakarla, Liangyu Zhao 等SIGCOMM 2024 · 被引用 43 次
相关 Paper
- ScheMoE: An Extensible Mixture-of-Experts Distributed Training System with Tasks SchedulingShaohuai Shi, Xinglin Pan, Qiang Wang, Chengjian Liu 等EuroSys 2024 · 被引用 26 次
- Efficient all-to-all Collective Communication Schedules for Direct-connect TopologiesPrithwish Basu, Liangyu Zhao, Jason Fantl, Siddharth Pal 等HPDC 2024 · 被引用 7 次
- FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts TrainingYunqi Gao, Bing Hu, Mahdi Boloursaz Mashhadi, A-Long Jin 等NeurIPS 2025 · 被引用 6 次
- SyCCL: Exploiting Symmetry for Efficient Collective Communication SchedulingJiamin Cao, Shangfeng Shi, Jiaqi Gao, Weisen Liu 等SIGCOMM 2025 · 被引用 15 次
- UBEP: Re-architecting Expert Parallelism Communication Library for Production SuperpodsYipeng Liu, Chang Liu, Si Shen, Jiaqi Zheng 等SIGCOMM 2026
