Towards Domain-Specific Network Transport for Distributed DNN Training
Hao Wang, Han Tian, Jingrong Chen, Xinchen Wan, Jiacheng Xia, Gaoxiong Zeng, Wei Bai, Junchen Jiang, Yong Wang, Kai Chen
摘要
Machine learning (ML) applications present rich characteristics to underlying network transport, yet little work has been done so far to systematically exploit these properties in transport design. This paper takes the initiative to pursue a domain-specific network transport, called MLT 1 , for distributed DNN training that fully embraces several unique characteristics of machine learning.
At its heart, MLT employs three simple-yet-effective techniques to form a 3-step progressive scheme against long tail latency caused by transient packet drops and queueing. First, it leverages the independencies among gradient updates to enable per-packet load balancing to minimize network hotspots without worrying about packet re-ordering. Then, if hotspot arises, it performs priority queueing/dropping based on the layers and magnitudes of gradients to optimize the model convergence. Lastly, if drop occurs, it enables bounded-loss tolerance-a certain amount of gradient losses tolerated by the DNN training without affecting the model accuracy. MLT is readily deployable with commodity switches and imposes minimal modifications on various DNN training libraries (e.g., TensorFlow, MXNet and PyTorch) and communication routines (e.g., PS and Ring All-reduce). We show, via both testbed experiments and simulations, that MLT effectively optimizes network tail latency and delivers up to 62.2% better end-to-end training performance over prior work.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- White-Boxing RDMA with Packet-Granular Software ControlChenxingyu Zhao, Jaehong Min, Ming Liu, Arvind KrishnamurthyNSDI 2025 · 被引用 28 次
- Design and Operation of Shared Machine Learning Clusters on CampusKaiqiang Xu, Decang Sun, Hao Wang, Zhenghang Ren 等ASPLOS 2025 · 被引用 24 次
- eTran: Extensible Kernel Transport with eBPFZhongjie Chen, Qingkai Meng, ChonLam Lao, Yifan Liu 等NSDI 2025 · 被引用 17 次
- Accelerating Neural Recommendation Training with Embedding SchedulingChaoliang Zeng, Xudong Liao, Xiaodian Cheng, Han Tian 等NSDI 2024 · 被引用 16 次
- ResCCL: Resource-Efficient Scheduling for Collective CommunicationTongrui Liu, Chenyang Hei, Fuliang Li, Chengxi Gao 等SIGCOMM 2025 · 被引用 11 次
它引用的顶会 Paper9
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi 等OSDI 2020 · 被引用 390 次
- ATP: In-network Aggregation for Multi-tenant LearningChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen 等NSDI 2021 · 被引用 359 次
- Aeolus: A Building Block for Proactive Transport in DatacentersShuihai Hu, Wei Bai, Gaoxiong Zeng, Zilong Wang 等SIGCOMM 2020 · 被引用 138 次
- Efficient sparse collective communication and its application to accelerate distributed deep learningJiawei Fei, Chen-Yu Ho, Atal Narayan Sahu, Marco Canini 等SIGCOMM 2021 · 被引用 120 次
相关 Paper
- OptiReduce: Resilient and Tail-Optimal AllReduce for Distributed Deep Learning in the CloudErtza Warraich, Omer Shabtai, Khalid Manaa, Shay Vargaftik 等NSDI 2025
- SiP-ML: high-bandwidth optical network interconnects for machine learning trainingMehrdad Khani Shirkoohi, Manya Ghobadi, Mohammad Alizadeh, Ziyi Zhu 等SIGCOMM 2021 · 被引用 94 次
- Towards timeout-less transport in commodity datacenter networksHwijoon Lim, Wei Bai, Yibo Zhu, Youngmok Jung 等EuroSys 2021 · 被引用 28 次
- Scaling Distributed Machine Learning with In-Network AggregationAmedeo Sapio, Marco Canini, Chen-Yu Ho, Jacob Nelson 等NSDI 2021
- Preemptive All-reduce Scheduling for Expediting Distributed DNN TrainingYixin Bao, Yanghua Peng, Yangrui Chen, Chuan WuINFOCOM 2020 · 被引用 67 次
