Mercury: A Simple Transport Layer Scheduler to Accelerate Distributed DNN Training
Qingyang Duan, Zeqin Wang, Yuedong Xu, Shaoteng Liu, Jun Wu
Abstract
Communication scheduling is crucial to improve the efficiency of training large deep learning models with data parallelism, in which the transmission order of layer-wise deep neural network (DNN) tensors is determined for a better computation-communication overlap. Prior approaches adopt tensor partitioning to enhance the priority scheduling with finer granularity. However, a startup time slot inserted before each tensor partition will neutralize this scheduling gain. Tuning the optimal partition size is difficult and the application-layer solutions cannot eliminate the partitioning overhead. In this paper, we propose Mercury, a simple transport layer scheduler that does not partition the tensors, but moves the priority scheduling to the transport layer at the packet granularity. The packets with the highest priority in the Mercury buffer will be transmitted first. Mercury achieves the near-optimal overlapping between communication and computation. It leverages immediate aggregation at the transport layer to enable the coincident gradient push and parameter pull. We implement Mercury in MXNet and conduct comprehensive experiments on five DNN models in an 8-node cluster with 10Gbps Ethernet. Experimental results show that Mercury can achieve about 1.18 2.18 × speedup over vanilla MXNet, and 1.08 2.04× speedup over the state-of-the-art tensor partitioning solution.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b99ef2c4-7880-4ac2-be31-b8ef9f8c337aRelated papers
- Preemptive All-reduce Scheduling for Expediting Distributed DNN TrainingYixin Bao, Yanghua Peng, Yangrui Chen, Chuan WuINFOCOM 2020 · 67 citations
- Libra: Contention-Aware GPU Thread Allocation for Data Parallel Training in High Speed NetworksYunzhuo Liu, Bo Jiang, Shizhen Zhao, Tao Lin et al.INFOCOM 2023 · 4 citations
- Exploiting Simultaneous Communications to Accelerate Data Parallel Distributed Deep LearningShaohuai Shi, Xiaowen Chu, Bo LiINFOCOM 2021 · 36 citations
- Concerto: Automatic Communication Optimization and Scheduling for Large-Scale Deep LearningShenggan Cheng, Shengjie Lin, Lansong Diao, Hao Wu et al.ASPLOS 2025 · 6 citations
- LEVELLER: Fair Communication Scheduling via Progress-Rate Awareness in Multi-Tenant Training ClustersGeng Li, Yang Li, Mingyuan Zang, Jie WuSIGCOMM 2026
