GADGET: Online Resource Optimization for Scheduling Ring-All-Reduce Learning Jobs
Menglu Yu, Ye Tian, Bo Ji, Chuan Wu, Hridesh Rajan, Jia Liu
Abstract
Fueled by advances in distributed deep learning (DDL), recent years have witnessed a rapidly growing demand for resource-intensive distributed/parallel computing to process DDL computing jobs. To resolve network communication bottleneck and load balancing issues in distributed computing, the so-called "ring-all-reduce" decentralized architecture has been increasingly adopted to remove the need for dedicated parameter servers. To date, however, there remains a lack of theoretical understanding on how to design resource optimization algorithms for efficiently scheduling ring-all-reduce DDL jobs in computing clusters. This motivates us to fill this gap by proposing a series of new resource scheduling designs for ring-all-reduce DDL jobs. Our contributions in this paper are threefold: i) We propose a new resource scheduling analytical model for ring-all-reduce deep learning, which covers a wide range of objectives in DDL performance optimization (e.g., excessive training avoidance, energy efficiency, fairness); ii) Based on the proposed performance analytical model, we develop an efficient resource scheduling algorithm called GADGET (greedy ring-all-reduce distributed graph embedding technique), which enjoys a provable strong performance guarantee; iii) We conduct extensive trace-driven experiments to demonstrate the effectiveness of the GADGET approach and its superiority over the state of the art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 542872bd-8f4d-456e-afed-25b97f30fef2Cited by top-tier papers1
Ask how each one uses itBuilds on2
- Preemptive All-reduce Scheduling for Expediting Distributed DNN TrainingYixin Bao, Yanghua Peng, Yangrui Chen, Chuan WuINFOCOM 2020 · 67 citations
- Themis: Fair and Efficient GPU Cluster SchedulingKshiteej Mahajan, Arjun Balasubramanian, Arjun Singhvi, Shivaram Venkataraman et al.NSDI 2020 · 22 citations
Related papers
- A Sum-of-Ratios Multi-Dimensional-Knapsack Decomposition for DNN Resource SchedulingMenglu Yu, Chuan Wu, Bo Ji, Jia LiuINFOCOM 2021 · 5 citations
- Communication Algorithm-Architecture Co-Design for Distributed Deep LearningJiayi Huang, Pritam Majumder, Sungkeun Kim, Abdullah Muzahid et al.ISCA 2021 · 44 citations
- Hare: Exploiting Inter-job and Intra-job Parallelism of Distributed Machine Learning on Heterogeneous GPUsFahao Chen, Peng Li, Celimuge Wu, Song GuoHPDC 2022 · 10 citations
- Multi-resource interleaving for deep learning trainingYihao Zhao, Yuanqiang Liu, Yanghua Peng, Yibo Zhu et al.SIGCOMM 2022 · 78 citations
- Accelerating Neural Recommendation Training with Embedding SchedulingChaoliang Zeng, Xudong Liao, Xiaodian Cheng, Han Tian et al.NSDI 2024 · 16 citations
