Multi-resource interleaving for deep learning training
Yihao Zhao, Yuanqiang Liu, Yanghua Peng, Yibo Zhu, Xuanzhe Liu, Xin Jin
Abstract
Training Deep Learning (DL) model requires multiple resource types, including CPUs, GPUs, storage IO, and network IO. Advancements in DL have produced a wide spectrum of models that have diverse usage patterns on different resource types. Existing DL schedulers focus on only GPU allocation, while missing the opportunity of packing jobs along multiple resource types.
We present Muri, a multi-resource cluster scheduler for DL workloads. Muri exploits multi-resource interleaving of DL training jobs to achieve high resource utilization and reduce job completion time (JCT). DL jobs have a unique staged, iterative computation pattern. In contrast to multi-resource schedulers for big data workloads that pack jobs in the space dimension, Muri leverages this unique pattern to interleave jobs on the same set of resources in the time dimension. Muri adapts Blossom algorithm to find the perfect grouping plan for single-GPU jobs with two resource types, and generalizes the algorithm to handle multi-GPU jobs with more than two types. We build a prototype of Muri and integrate it with PyTorch. Experiments on a cluster with 64 GPUs demonstrate that Muri improves the average JCT by up to 3.6× and the makespan by up to 1.6× over existing DL schedulers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af3aa20c-e081-40fc-9f17-151551a4fa59Cited by top-tier papers15
- AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning ServingZhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu et al.OSDI 2023 · 211 citations
- CASSINI: Network-Aware Job Scheduling in Machine Learning ClustersSudarsanan Rajasekaran, Manya Ghobadi, Aditya AkellaNSDI 2024 · 144 citations
- Beware of Fragmentation: Scheduling GPU-Sharing Workloads with Fragmentation Gradient DescentQizhen Weng, Lingyun Yang, Yinghao Yu, Wei Wang et al.USENIX ATC 2023 · 115 citations
- ElasticFlow: An Elastic Serverless Training Platform for Distributed Deep LearningDiandian Gu, Yihao Zhao, Yinmin Zhong, Yifan Xiong et al.ASPLOS 2023 · 62 citations
- Crux: GPU-Efficient Communication Scheduling for Deep Learning TrainingJiamin Cao, Yu Guan, Kun Qian, Jiaqi Gao et al.SIGCOMM 2024 · 60 citations
Builds on11
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi et al.OSDI 2020 · 390 citations
- Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning WorkloadsDeepak Narayanan, Keshav Santhanam, Fiodar Kazhamiaka, Amar Phanishayee et al.OSDI 2020 · 286 citations
- AntMan: Dynamic Scaling on GPU Clusters for Deep LearningWencong Xiao, Shiru Ren, Yong Li, Yang Zhang et al.OSDI 2020 · 260 citations
- Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep LearningAurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger et al.OSDI 2021 · 258 citations
- Analyzing and Mitigating Data Stalls in DNN TrainingJayashree Mohan, Amar Phanishayee, Ashish Raniwala, Vijay ChidambaramVLDB 2021 · 142 citations
Related papers
- Looking Beyond GPUs for DNN Scheduling on Multi-Tenant ClustersJayashree Mohan, Amar Phanishayee, Janardhan Kulkarni, Vijay ChidambaramOSDI 2022 · 91 citations
- SiloD: A Co-design of Caching and Scheduling for Deep Learning ClustersHanyu Zhao, Zhenhua Han, Zhi Yang, Quanlu Zhang et al.EuroSys 2023 · 22 citations
- A Sum-of-Ratios Multi-Dimensional-Knapsack Decomposition for DNN Resource SchedulingMenglu Yu, Chuan Wu, Bo Ji, Jia LiuINFOCOM 2021 · 5 citations
- An efficient and non-intrusive GPU scheduling framework for deep learning training systemsShaoqi Wang, Oscar J. Gonzalez, Xiaobo Zhou, Thomas Williams et al.SC 2020 · 21 citations
- PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU ClustersRutwik Jain, Brandon Tran, Keting Chen, Matthew D. Sinclair et al.SC 2024 · 10 citations
