Understanding and Optimizing Workloads for Unified Resource Management in Large Cloud Platforms
Chengzhi Lu, Huanle Xu, Kejiang Ye, Guoyao Xu, Liping Zhang, Guodong Yang, Chengzhong Xu
摘要
To fully utilize computing resources, cloud providers such as Google and Alibaba choose to co-locate online services with batch processing applications in their data centers. By implementing unified resource management policies, different types of complex computing jobs request resources in a consistent way, which can help data centers achieve global optimal scheduling and provide computing power with higher quality. To understand this new scheduling paradigm, in this paper, we first present an in-depth study of Alibaba's unified scheduling workloads. Our study focuses on the characterization of resource utilization, the application running performance, and scheduling scalability. We observe that although computing resources are significantly over-committed under unified scheduling, the resource utilization in Alibaba data centers is still low. In addition, existing resource usage predictors tend to make severe overestimations. At the same time, tasks within the same application behave fairly consistently, and the running performance of tasks can be well-profiled with respect to resource contention on the corresponding physical host.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- EXIST: Enabling Extremely Efficient Intra-Service Tracing Observability in DatacentersXinkai Wang, Xiaofeng Hou, Chao Li, Yuancheng Li 等ASPLOS 2025 · 被引用 4 次
- AUM: Unleashing the Efficiency Potential of Shared Processors with Accelerator Units for LLM ServingXinkai Wang, Chao Li, Yiming Zhuansun, Jinyang Guo 等HPCA 2026 · 被引用 2 次
- MerKury: Adaptive Resource Allocation to Enhance the Kubernetes Performance for Large-Scale ClustersJiayin Luo, Xinkui Zhao, Yuxin Ma, Shengye Pang 等WWW 2025
相关 Paper
- Take it to the limit: peak prediction-driven resource overcommitment in datacentersNoman Bashir, Nan Deng, Krzysztof Rzadca, David Irwin 等EuroSys 2021 · 被引用 60 次
- Rhythm: component-distinguishable workload deployment in datacentersLaiping Zhao, Yanan Yang, Kaixuan Zhang, Xiaobo Zhou 等EuroSys 2020 · 被引用 49 次
- MLaaS in the Wild: Workload Analysis and Scheduling in Large-Scale Heterogeneous GPU ClustersQizhen Weng, Wencong Xiao, Yinghao Yu, Wei Wang 等NSDI 2022
- Scheduling Cloud Block Storage Proactively and Reactively with OmarXinqi Chen, Weidong Zhang, Zhongyu Wang, Erci Xu 等EuroSys 2026
- Eva: Cost-Efficient Cloud-Based Cluster SchedulingTzu-Tao Chang, Shivaram VenkataramanEuroSys 2025 · 被引用 2 次
