Take it to the limit: peak prediction-driven resource overcommitment in datacenters
Noman Bashir, Nan Deng, Krzysztof Rzadca, David Irwin, Sree Kodak, Rohit Jnagal
摘要
To increase utilization, datacenter schedulers often overcommit resources where the sum of resources allocated to the tasks on a machine exceeds its physical capacity. Setting the right level of overcommitment is a challenging problem: low overcommitment leads to wasted resources, while high overcommitment leads to task performance degradation. In this paper, we take a first principles approach to designing and evaluating overcommit policies by asking a basic question: assuming complete knowledge of each task's future resource usage, what is the safest overcommit policy that yields the highest utilization? We call this policy the peak oracle. We then devise practical overcommit policies that mimic this peak oracle by predicting future machine resource usage. We simulate our overcommit policies using the recently-released Google cluster trace, and show that they result in higher utilization and less overcommit errors than policies based on per-task allocations. We also deploy these policies to machines inside Google's datacenters serving its internal production workload. We show that our overcommit policies increase these machines' usable CPU capacity by 10-16% compared to no overcommitment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- On the Limitations of Carbon-Aware Temporal and Spatial Workload Shifting in the CloudThanathorn Sukprasert, Abel Souza, Noman Bashir, David Irwin 等EuroSys 2024 · 被引用 73 次
- AUDIBLE: A Convolution-Based Resource Allocator for Oversubscribing Burstable Virtual MachinesSeyed Ali Jokar Jandaghi, Kaveh Mahdaviani, Amirhossein Mirhosseini, Sameh Elnikety 等ASPLOS 2024 · 被引用 1 次
- Scheduling Cloud Block Storage Proactively and Reactively with OmarXinqi Chen, Weidong Zhang, Zhongyu Wang, Erci Xu 等EuroSys 2026
- MDK: Rethinking the Data Center Memory Reclamation ProblemShaurya Patel, Suli Yang, Yawen Wang, Kan Wu 等OSDI 2026
它引用的顶会 Paper5
- FIRM: An Intelligent Fine-grained Resource Management Framework for SLO-Oriented MicroservicesHaoran Qiu, Subho S. Banerjee, Saurabh Jha, Zbigniew T. Kalbarczyk 等OSDI 2020 · 被引用 350 次
- Providing SLOs for Resource-Harvesting VMs in Cloud PlatformsPradeep Ambati, Iñigo Goiri, Felipe Vieira Frujeri, Alper Gun 等OSDI 2020 · 被引用 101 次
- Rhythm: component-distinguishable workload deployment in datacentersLaiping Zhao, Yanan Yang, Kaixuan Zhang, Xiaobo Zhou 等EuroSys 2020 · 被引用 49 次
- Seagull: An Infrastructure for Load Prediction and Optimized Resource AllocationOlga Poppe, Tayo Amuneke, Dalitso Banda, Aritra De 等VLDB 2021 · 被引用 37 次
- Improving resource utilization by timely fine-grained schedulingTatiana Jin, Zhenkun Cai, Boyang Li, Chengguang Zheng 等EuroSys 2020 · 被引用 27 次
相关 Paper
- Understanding and Optimizing Workloads for Unified Resource Management in Large Cloud PlatformsChengzhi Lu, Huanle Xu, Kejiang Ye, Guoyao Xu 等EuroSys 2023 · 被引用 34 次
- Preventing Network Bottlenecks: Accelerating Datacenter Services with Hotspot-Aware Placement for Compute and StorageHamid Hajabdolali Bazzaz, Yingjie Bi, Weiwu Pang, Minlan Yu 等NSDI 2025 · 被引用 3 次
- Learning Cooperative Oversubscription for Cloud by Chance-Constrained Multi-Agent Reinforcement LearningJunjie Sheng, Lu Wang, Fangkai Yang, Bo Qiao 等WWW 2023 · 被引用 10 次
- Autopilot: workload autoscaling at GoogleKrzysztof Rzadca, Pawel Findeisen, Jacek Swiderski, Przemyslaw Zych 等EuroSys 2020 · 被引用 299 次
- Prediction-Based Power Oversubscription in Cloud PlatformsAlok Gautam Kumbhare, Reza Azimi, Ioannis Manousakis, Anand Bonde 等USENIX ATC 2021 · 被引用 90 次
