MLaaS in the Wild: Workload Analysis and Scheduling in Large-Scale Heterogeneous GPU Clusters
Qizhen Weng, Wencong Xiao, Yinghao Yu, Wei Wang, Cheng Wang, Jian He, Yong Li, Liping Zhang, Wei Lin, Yu Ding
摘要
With the sustained technological advances in machine learning (ML) and the availability of massive datasets recently, tech companies are deploying large ML-as-a-Service (MLaaS) clouds, often with heterogeneous GPUs, to provision a host of ML applications. However, running diverse ML workloads in heterogeneous GPU clusters raises a number of challenges. In this paper, we present a characterization study of a two-month workload trace collected from a production MLaaS cluster with over 6,000 GPUs in Alibaba. We explain the challenges posed to cluster scheduling, including the low GPU utilization, the long queueing delays, the presence of hard-to-schedule tasks demanding high-end GPUs with picky scheduling requirements, the imbalance load across heterogeneous machines, and the potential bottleneck on CPUs. We describe our current solutions and call for further investigations into the challenges that remain open to address. We have released the trace for public access, which is the most comprehensive in terms of the workloads and cluster scale.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper96
- Zeus: Understanding and Optimizing GPU Energy Consumption of DNN TrainingJie You, Jae-Won Chung, Mosharaf ChowdhuryNSDI 2023 · 被引用 220 次
- AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning ServingZhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu 等OSDI 2023 · 被引用 211 次
- Characterization of Large Language Model Development in the DatacenterQinghao Hu, Zhisheng Ye, Zerui Wang, Guoteng Wang 等NSDI 2024 · 被引用 192 次
- Beware of Fragmentation: Scheduling GPU-Sharing Workloads with Fragmentation Gradient DescentQizhen Weng, Lingyun Yang, Yinghao Yu, Wei Wang 等USENIX ATC 2023 · 被引用 115 次
- Transparent GPU Sharing in Container Clouds for Deep Learning WorkloadsBingyang Wu, Zili Zhang, Zhihao Bai, Xuanzhe Liu 等NSDI 2023 · 被引用 112 次
它引用的顶会 Paper9
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- MLPerf Inference BenchmarkVijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson 等ISCA 2020 · 被引用 517 次
- Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning WorkloadsDeepak Narayanan, Keshav Santhanam, Fiodar Kazhamiaka, Amar Phanishayee 等OSDI 2020 · 被引用 286 次
- Balancing efficiency and fairness in heterogeneous GPU clusters for deep learningShubham Chaudhary, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra 等EuroSys 2020 · 被引用 135 次
- Metis: learning to schedule long-running applications in shared container clusters at scaleLuping Wang, Qizhen Weng, Wei Wang, Chen Chen 等SC 2020 · 被引用 46 次
相关 Paper
- AntMan: Dynamic Scaling on GPU Clusters for Deep LearningWencong Xiao, Shiru Ren, Yong Li, Yang Zhang 等OSDI 2020 · 被引用 260 次
- Hare: Exploiting Inter-job and Intra-job Parallelism of Distributed Machine Learning on Heterogeneous GPUsFahao Chen, Peng Li, Celimuge Wu, Song GuoHPDC 2022 · 被引用 10 次
- PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU ClustersRutwik Jain, Brandon Tran, Keting Chen, Matthew D. Sinclair 等SC 2024 · 被引用 10 次
- Heet: Accelerating Elastic Training in Heterogeneous Deep Learning ClustersZizhao Mo, Huanle Xu, Chengzhong XuASPLOS 2024 · 被引用 22 次
- Understanding and Optimizing Workloads for Unified Resource Management in Large Cloud PlatformsChengzhi Lu, Huanle Xu, Kejiang Ye, Guoyao Xu 等EuroSys 2023 · 被引用 34 次
