Take it to the limit: peak prediction-driven resource overcommitment in datacenters
Noman Bashir, Nan Deng, Krzysztof Rzadca, David Irwin, Sree Kodak, Rohit Jnagal
Abstract
To increase utilization, datacenter schedulers often overcommit resources where the sum of resources allocated to the tasks on a machine exceeds its physical capacity. Setting the right level of overcommitment is a challenging problem: low overcommitment leads to wasted resources, while high overcommitment leads to task performance degradation. In this paper, we take a first principles approach to designing and evaluating overcommit policies by asking a basic question: assuming complete knowledge of each task's future resource usage, what is the safest overcommit policy that yields the highest utilization? We call this policy the peak oracle. We then devise practical overcommit policies that mimic this peak oracle by predicting future machine resource usage. We simulate our overcommit policies using the recently-released Google cluster trace, and show that they result in higher utilization and less overcommit errors than policies based on per-task allocations. We also deploy these policies to machines inside Google's datacenters serving its internal production workload. We show that our overcommit policies increase these machines' usable CPU capacity by 10-16% compared to no overcommitment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- On the Limitations of Carbon-Aware Temporal and Spatial Workload Shifting in the CloudThanathorn Sukprasert, Abel Souza, Noman Bashir, David Irwin et al.EuroSys 2024 · 73 citations
- AUDIBLE: A Convolution-Based Resource Allocator for Oversubscribing Burstable Virtual MachinesSeyed Ali Jokar Jandaghi, Kaveh Mahdaviani, Amirhossein Mirhosseini, Sameh Elnikety et al.ASPLOS 2024 · 1 citation
- Scheduling Cloud Block Storage Proactively and Reactively with OmarXinqi Chen, Weidong Zhang, Zhongyu Wang, Erci Xu et al.EuroSys 2026
- MDK: Rethinking the Data Center Memory Reclamation ProblemShaurya Patel, Suli Yang, Yawen Wang, Kan Wu et al.OSDI 2026
Builds on5
- FIRM: An Intelligent Fine-grained Resource Management Framework for SLO-Oriented MicroservicesHaoran Qiu, Subho S. Banerjee, Saurabh Jha, Zbigniew T. Kalbarczyk et al.OSDI 2020 · 350 citations
- Providing SLOs for Resource-Harvesting VMs in Cloud PlatformsPradeep Ambati, Iñigo Goiri, Felipe Vieira Frujeri, Alper Gun et al.OSDI 2020 · 101 citations
- Rhythm: component-distinguishable workload deployment in datacentersLaiping Zhao, Yanan Yang, Kaixuan Zhang, Xiaobo Zhou et al.EuroSys 2020 · 49 citations
- Seagull: An Infrastructure for Load Prediction and Optimized Resource AllocationOlga Poppe, Tayo Amuneke, Dalitso Banda, Aritra De et al.VLDB 2021 · 37 citations
- Improving resource utilization by timely fine-grained schedulingTatiana Jin, Zhenkun Cai, Boyang Li, Chengguang Zheng et al.EuroSys 2020 · 27 citations
Related papers
- Understanding and Optimizing Workloads for Unified Resource Management in Large Cloud PlatformsChengzhi Lu, Huanle Xu, Kejiang Ye, Guoyao Xu et al.EuroSys 2023 · 34 citations
- Preventing Network Bottlenecks: Accelerating Datacenter Services with Hotspot-Aware Placement for Compute and StorageHamid Hajabdolali Bazzaz, Yingjie Bi, Weiwu Pang, Minlan Yu et al.NSDI 2025 · 3 citations
- Learning Cooperative Oversubscription for Cloud by Chance-Constrained Multi-Agent Reinforcement LearningJunjie Sheng, Lu Wang, Fangkai Yang, Bo Qiao et al.WWW 2023 · 10 citations
- Autopilot: workload autoscaling at GoogleKrzysztof Rzadca, Pawel Findeisen, Jacek Swiderski, Przemyslaw Zych et al.EuroSys 2020 · 299 citations
- Prediction-Based Power Oversubscription in Cloud PlatformsAlok Gautam Kumbhare, Reza Azimi, Ioannis Manousakis, Anand Bonde et al.USENIX ATC 2021 · 90 citations
