USENIX ATC2024顶会
Starburst: A Cost-aware Scheduler for Hybrid Cloud
Michael Luo, Siyuan Zhuang, Suryaprakash Vengadesan, Romil Bhardwaj, Justin Chang, Eric J. Friedman, Scott Shenker, Ion Stoica
摘要
To efficiently tackle bursts in job demand, organizations employ hybrid cloud architectures to scale their batch workloads from their private clusters to public cloud. This requires transforming cluster schedulers into cloud-enabled versions to navigate the tradeoff between cloud costs and scheduler objectives such as job completion time (JCT). However, our analysis over production-level traces show that existing cloud-enabled schedulers incur inefficient cost-JCT trade-offs due to low cluster utilization.
We present Starburst, a system that maximizes cluster utilization to streamline the cost-JCT tradeoff. Starburst's scheduler dynamically controls jobs' waiting times to improve utilization-it assigns longer waits for large jobs to increase their chances of running on the cluster, and shorter waits to small jobs to increase their chances of running on the cloud. To offer configurability, Starburst provides system administrators a simple waiting budget framework to tune their position on the cost-JCT curve. A departure from traditional cluster schedulers, Starburst operates as a higher-level resource manager over a private cluster and dynamic cloud clusters. Simulations over production-level traces and real-world experiments on a 32-GPU private cluster show that Starburst can reduce cloud costs by up to 54-91% over existing cluster managers, while increasing average JCT by at most 5.8%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- FIRM: An Intelligent Fine-grained Resource Management Framework for SLO-Oriented MicroservicesHaoran Qiu, Subho S. Banerjee, Saurabh Jha, Zbigniew T. Kalbarczyk 等OSDI 2020 · 被引用 350 次
- Autopilot: workload autoscaling at GoogleKrzysztof Rzadca, Pawel Findeisen, Jacek Swiderski, Przemyslaw Zych 等EuroSys 2020 · 被引用 299 次
- Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning WorkloadsDeepak Narayanan, Keshav Santhanam, Fiodar Kazhamiaka, Amar Phanishayee 等OSDI 2020 · 被引用 286 次
- Sinan: ML-based and QoS-aware resource management for cloud microservicesYanqi Zhang, Weizhe Hua, Zhuangzhuang Zhou, G. Edward Suh 等ASPLOS 2021 · 被引用 226 次
- AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning ServingZhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu 等OSDI 2023 · 被引用 211 次
相关 Paper
- Waiting game: optimally provisioning fixed resources for cloud-enabled schedulersPradeep Ambati, Noman Bashir, David Irwin, Prashant J. ShenoySC 2020 · 被引用 14 次
- Eva: Cost-Efficient Cloud-Based Cluster SchedulingTzu-Tao Chang, Shivaram VenkataramanEuroSys 2025 · 被引用 2 次
- Qonductor: A Cloud Orchestrator for Quantum ComputingEmmanouil Giortamis, Francisco Romão, Nathaniel Tornow, Dmitry Lugovoy 等SC 2025 · 被引用 4 次
- Going Green for Less Green: Optimizing the Cost of Reducing Cloud Carbon EmissionsWalid A. Hanafy, Qianlin Liang, Noman Bashir, Abel Souza 等ASPLOS 2024 · 被引用 43 次
- Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep LearningAurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger 等OSDI 2021 · 被引用 258 次
