Learning Cooperative Oversubscription for Cloud by Chance-Constrained Multi-Agent Reinforcement Learning
Junjie Sheng, Lu Wang, Fangkai Yang, Bo Qiao, Hang Dong, Xiangfeng Wang, Bo Jin, Jun Wang, Si Qin, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang
Abstract
Oversubscription is a common practice for improving cloud resource utilization. It allows the cloud service provider to sell more resources than the physical limit, assuming not all users would fully utilize the resources simultaneously. However, how to design an oversubscription policy that improves utilization while satisfying the some safety constraints remains an open problem. Existing methods and industrial practices are over-conservative, ignoring the coordination of diverse resource usage patterns and probabilistic constraints. To address these two limitations, this paper formulates the oversubscription for cloud as a chance-constrained optimization problem and propose an effective Chance-Constrained Multi-Agent Reinforcement Learning (C2MARL) method to solve this problem. Specifically, C2MARL reduces the number of constraints by considering their upper bounds and leverages a multi-agent reinforcement learning paradigm to learn a safe and optimal coordination policy. We evaluate our C2MARL on an internal cloud platform and public cloud datasets. Experiments show that our C2MARL outperforms existing methods in improving utilization (20% ∼ 86%) under different levels of safety constraints.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f81ee440-a14e-4a0b-9c1a-9b6744904766Cited by top-tier papers2
- Coach: Exploiting Temporal Patterns for All-Resource Oversubscription in Cloud PlatformsBenjamin Reidys, Pantea Zardoshti, Íñigo Goiri, Celine Irvene et al.ASPLOS 2025 · 9 citations
- On the Hardness of Constrained Cooperative Multi-Agent Reinforcement LearningZiyi Chen, Yi Zhou, Heng HuangICLR 2024 · 6 citations
Builds on3
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud ProviderMohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Irfan Chaudhry et al.USENIX ATC 2020 · 946 citations
- Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision ProcessesDongsheng Ding, Kaiqing Zhang, Tamer Basar, Mihailo R. JovanovicNeurIPS 2020 · 252 citations
- Protean: VM Allocation Service at ScaleOri Hadary, Luke Marshall, Ishai Menache, Abhisek Pan et al.OSDI 2020 · 189 citations
Related papers
- AWARE: Automate Workload Autoscaling with Reinforcement Learning in Production Cloud SystemsHaoran Qiu, Weichao Mao, Chen Wang, Hubertus Franke et al.USENIX ATC 2023 · 95 citations
- Following the Usage, Not the Request: Risk-Aware Task Scheduling with Overbooking in Edge CloudsTie Ma, Shan Zhang, Xiaoyu Zhang, Zichuan Zheng et al.INFOCOM 2026
- Take it to the limit: peak prediction-driven resource overcommitment in datacentersNoman Bashir, Nan Deng, Krzysztof Rzadca, David Irwin et al.EuroSys 2021 · 60 citations
- Flipping-based Policy for Chance-Constrained Markov Decision ProcessesXun Shen, Shuo Jiang, Akifumi Wachi, Kazumune Hashimoto et al.NeurIPS 2024 · 5 citations
- Multi-Unit Auctions for Allocating Chance-Constrained ResourcesAnna Gautier, Bruno Lacerda, Nick Hawes, Michael J. WooldridgeAAAI 2023 · 5 citations
