FAST: An Efficient Scheduler for All-to-All GPU Communication
Yiran Lei, Dongjoo Lee, Liangyu Zhao, Daniar Kurniawan, Chanmyeong Kim, Heetaek Jeong, Changsu Kim, Hyeonseong Choi, Liangcheng Yu, Arvind Krishnamurthy, Justine Sherry, Eriko Nurvitadhi
Abstract
All-to-All(v) communication is a critical primitive in modern machine learning workloads, particularly mixture-of-experts (MoE) models. Unfortunately, efficient scheduling is challenging due to workload skew, heterogeneous two-tier fabrics, and incast congestion, compounded by the dynamic nature of MoE workloads, where traffic shifts every few hundred milliseconds. Existing schedulers are hardly scalable, incurring seconds to hours of synthesis time, making them impractical. We present FAST, an efficient All-to-All(v) scheduler. FAST addresses skew through intra-server rebalancing and enforces balanced, one-to-one scale-out transfers that avoid incast. Evaluated extensively on both NVIDIA H200 and AMD MI300X clusters, FAST consistently outperforms state-of-the-art solutions on skewed workloads while reducing synthesis time by orders of magnitude.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 66beffd4-2ec4-4904-9b16-ef29c788d020Cited by top-tier papers3
- TEMP: A Memory Efficient Physical-Aware Tensor Partition-Mapping Framework on Wafer-Scale ChipsHuizheng Wang, Taiquan Wei, Zichuan Wang, Dingcheng Jiang et al.HPCA 2026 · 2 citations
- StreamEP: Straggler-Tolerant MoE Decoding without Communication BarriersYizhuo Liang, Shaoyu Wang, Jaeyong Song, Yanqi Zhou et al.SOSP 2026
- Understanding and Profiling the Accelerator Chiplet Network Using PingPointJunyeol Ryu, Ming Liu, Matthew D. SinclairSIGCOMM 2026
Builds on9
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleSamyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang et al.ICML 2022 · 523 citations
- Accelerating Distributed MoE Training and Inference with LinaJiamin Li, Yimin Jiang, Yibo Zhu, Cong Wang et al.USENIX ATC 2023 · 191 citations
- RDMA over Ethernet for Distributed Training at Meta ScaleAdithya Gangidi, Rui Miao, Shengbao Zheng, Sai Jayesh Bondu et al.SIGCOMM 2024 · 171 citations
- Synthesizing optimal collective algorithmsZixian Cai, Zhengyang Liu, Saeed Maleki, Madanlal Musuvathi et al.PPoPP 2021 · 64 citations
- Rethinking Machine Learning Collective Communication as a Multi-Commodity Flow ProblemXuting Liu, Behnaz Arzani, Siva Kesava Reddy Kakarla, Liangyu Zhao et al.SIGCOMM 2024 · 43 citations
Related papers
- ScheMoE: An Extensible Mixture-of-Experts Distributed Training System with Tasks SchedulingShaohuai Shi, Xinglin Pan, Qiang Wang, Chengjian Liu et al.EuroSys 2024 · 26 citations
- Efficient all-to-all Collective Communication Schedules for Direct-connect TopologiesPrithwish Basu, Liangyu Zhao, Jason Fantl, Siddharth Pal et al.HPDC 2024 · 7 citations
- FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts TrainingYunqi Gao, Bing Hu, Mahdi Boloursaz Mashhadi, A-Long Jin et al.NeurIPS 2025 · 6 citations
- SyCCL: Exploiting Symmetry for Efficient Collective Communication SchedulingJiamin Cao, Shangfeng Shi, Jiaqi Gao, Weisen Liu et al.SIGCOMM 2025 · 15 citations
- UBEP: Re-architecting Expert Parallelism Communication Library for Production SuperpodsYipeng Liu, Chang Liu, Si Shen, Jiaqi Zheng et al.SIGCOMM 2026
