Better Together: Jointly Optimizing ML Collective Scheduling and Execution Planning using SYNDICATE
Kshiteej Mahajan, Ching-Hsiang Chu, Srinivas Sridharan, Aditya Akella
Abstract
Emerging ML training deployments are trending towards larger models, and hybrid-parallel training that is not just dominated by compute-intensive all-reduce for gradient aggregation but also bandwidth-intensive collectives (e.g., all-to-all). These emerging collectives exacerbate the communication bottlenecks despite heterogeneous network interconnects with ample multipath opportunities. In this work, we propose SYNDICATE, a systematic, general framework to minimize communication bottlenecks and speed up training for both state-of-the-art and future large-scale models and interconnects. SYNDICATE proposes a novel abstraction, the motif, to break large communication work as smaller pieces as part of execution planning. SYNDICATE also does joint optimization of scheduling and execution planning by rethinking the interfaces in the networking systems stacks used for ML training. Motifs afford greater flexibility during scheduling and the joint optimizer exploits this flexibility by packing and ordering communication work so as to maximize both network utilization and overlap with compute. This improves the speed of training state-of-the-art large models by 21-74%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext df3bbfec-f0d7-4d81-8dc6-0a245281194aCited by top-tier papers16
- CASSINI: Network-Aware Job Scheduling in Machine Learning ClustersSudarsanan Rajasekaran, Manya Ghobadi, Aditya AkellaNSDI 2024 · 144 citations
- RLinf: Flexible and Efficient Large-Scale Reinforcement Learning via Macro-to-Micro Flow TransformationChao Yu, Yuanqing Wang, Zhen Guo, Hao Lin et al.OSDI 2026 · 63 citations
- Crux: GPU-Efficient Communication Scheduling for Deep Learning TrainingJiamin Cao, Yu Guan, Kun Qian, Jiaqi Gao et al.SIGCOMM 2024 · 60 citations
- Rethinking Machine Learning Collective Communication as a Multi-Commodity Flow ProblemXuting Liu, Behnaz Arzani, Siva Kesava Reddy Kakarla, Liangyu Zhao et al.SIGCOMM 2024 · 43 citations
- FlashMoE: Fast Distributed MoE in a Single KernelOsayamen Jonathan Aimuyo, Byungsoo Oh, Rachee SinghNeurIPS 2025 · 22 citations
Builds on3
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi et al.OSDI 2020 · 390 citations
Related papers
- Themis: a network bandwidth-aware collective scheduling policy for distributed training of DL modelsSaeed Rashidi, William Won, Sudarshan Srinivasan, Srinivas Sridharan et al.ISCA 2022 · 40 citations
- PipeComm: Maximizing Link Utilization Through Pipeline-Aware Collective Communication SynthesisRuifan Xu, Yuze Luo, Yuhao Meng, Size Zheng et al.ISCA 2026
- Preemptive All-reduce Scheduling for Expediting Distributed DNN TrainingYixin Bao, Yanghua Peng, Yangrui Chen, Chuan WuINFOCOM 2020 · 67 citations
- Centauri: Enabling Efficient Scheduling for Communication-Computation Overlap in Large Model Training via Communication PartitioningChang Chen, Xiuhong Li, Qianchao Zhu, Jiangfei Duan et al.ASPLOS 2024 · 52 citations
- SkipReduce: (Interconnection) Network Sparsity to Accelerate Distributed Machine LearningHans Kasan, Dennis Abts, Jungwook Choi, John KimMICRO 2025 · 1 citation
