SpiderFlow: Efficient Topology-Aware Scheduling for LLM Training Across Decentralized GPU Clusters
Zihan Chang, Shuibing He, Bo Zhou, Sheng Xiao, Siling Yang, Rui Wang, Zhe Pan
Abstract
In response to increasing demands for largescale machine learning training jobs, many organizations have deployed GPU clusters across geographically distributed regions. However, existing Integer Linear Programming (ILP)-or genetic-based cross-cluster training approaches largely overlook the topology of decentralized clusters, lacking both topology-aware task scheduling mechanisms and automated model parallelization strategies. As a result, naively applying these optimization-based methods in cross-cluster settings leads to prohibitive scheduling overhead, due to the drastically enlarged search space induced by complex inter-cluster topologies. To address these challenges, we propose SpiderFlow, a topologyaware scheduling system specifically designed for decentralized GPU clusters. We formulate cross-cluster task scheduling as a graph optimization problem and introduce SpinSearch, a low-overhead topology-aware scheduling algorithm. In addition, for automated model parallelization, we propose Topology-aware Parallelism Automation (TPA), a two-level scheduling framework that combines heuristic methods at the inter-cluster level with ILP-based optimization within clusters, effectively reducing the search space while maintaining high training throughput with substantially lower scheduling overhead. We evaluate SpiderFlow on a physical platform comprising 8 decentralized clusters, as well as on a simulation platform with up to 64 decentralized clusters. Experimental results demonstrate that SpiderFlow reduces job completion time (JCT) by 1.2-1.3×, improves throughput by 1.12-1.25×, and reduces scheduling overhead by 20-90× on average compared to state-of-the-art scheduling systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dfc7d28d-b103-4de4-9ce5-88cc7d084bc7Builds on8
- Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep LearningAurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger et al.OSDI 2021 · 258 citations
- Decentralized Training of Foundation Models in Heterogeneous EnvironmentsBinhang Yuan, Yongjun He, Jared Davis, Tianyi Zhang et al.NeurIPS 2022 · 157 citations
- Characterization and prediction of deep learning workloads in large-scale GPU datacentersQinghao Hu, Peng Sun, Shengen Yan, Yonggang Wen et al.SC 2021 · 136 citations
- Elastic Resource Sharing for Distributed Deep LearningChangho Hwang, Taehyun Kim, Sunghyun Kim, Jinwoo Shin et al.NSDI 2021 · 111 citations
- Looking Beyond GPUs for DNN Scheduling on Multi-Tenant ClustersJayashree Mohan, Amar Phanishayee, Janardhan Kulkarni, Vijay ChidambaramOSDI 2022 · 91 citations
Related papers
- Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-DesignChunyu Xue, Weihao Cui, Quan Chen, Chen Chen et al.EuroSys 2026
- Themis: Fair and Efficient GPU Cluster SchedulingKshiteej Mahajan, Arjun Balasubramanian, Arjun Singhvi, Shivaram Venkataraman et al.NSDI 2020 · 22 citations
- Sailor: Automating Distributed Training over Dynamic, Heterogeneous, and Geo-distributed ClustersFoteini Strati, Zhendong Zhang, George Manos, Ixeia Sánchez Périz et al.SOSP 2025 · 2 citations
- HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU ClustersAntian Liang, Zhigang Zhao, Kai Zhang, Xuri Shi et al.EuroSys 2026 · 1 citation
- HetAuto: Cross-Cluster Auto-Parallelism for Heterogeneous Distributed TrainingGuicheng Qi, Junwei Su, Liqi Yang, Tao Li et al.EuroSys 2026 · 1 citation
