PhOrch: Proactive Phase-Level Flow Path Orchestration For Contention-Free LLM Training
Ziyang Zou, Shuangwu Chen, Tao Zhang, Huihuang Qin, Jian Yang, Xiaobin Tan, Dong Jin
Abstract
The growing scale of large language models (LLMs) has made communication overhead a critical bottleneck in distributed training, primarily due to imbalanced traffic loads. Existing load balancing methods often lead to severe flow contention when handling low-entropy and high-volume LLM training flows. Motivated by the point-to-point pattern in each collective communication phase and the inherent periodicity of training traffic, we propose PhOrch, a proactive phase-level contention-free flow path orchestration framework tailored for LLM training workloads. We formulate the orchestration as an optimization problem, which is typically NP-hard. To tackle this problem, we develop a segmented edge coloring algorithm for bipartite multigraphs, which efficiently assigns flow paths while avoiding contention. Evaluation results demonstrate that PhOrch reduces the per-cycle communication time by 60% compared to the state-of-the-art methods and achieves contention-free training traffic in non-oversubscribed topologies, indicating that our method substantially mitigates the communication bottleneck during the LLM training process.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 47637eee-8fb0-4024-90e4-01a24242b69eRelated papers
- Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUsGuoliang He, Youhe Jiang, Wencong Xiao, Kaihua Jiang et al.NeurIPS 2025 · 10 citations
- Crux: GPU-Efficient Communication Scheduling for Deep Learning TrainingJiamin Cao, Yu Guan, Kun Qian, Jiaqi Gao et al.SIGCOMM 2024 · 60 citations
- GeoOrchestra: Orchestrating Heterogeneous Geo-Distributed Training with Network-Aware SchedulingTing Liu, Qinghua Wu, Jun Zhou, Yuan Sun et al.SIGCOMM 2026
- Alibaba HPN: A Data Center Network for Large Language Model TrainingKun Qian, Yongqing Xi, Jiamin Cao, Jiaqi Gao et al.SIGCOMM 2024 · 173 citations
- Better Together: Jointly Optimizing ML Collective Scheduling and Execution Planning using SYNDICATEKshiteej Mahajan, Ching-Hsiang Chu, Srinivas Sridharan, Aditya AkellaNSDI 2023 · 43 citations
