Learning Cooperative Oversubscription for Cloud by Chance-Constrained Multi-Agent Reinforcement Learning
Junjie Sheng, Lu Wang, Fangkai Yang, Bo Qiao, Hang Dong, Xiangfeng Wang, Bo Jin, Jun Wang, Si Qin, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang
摘要
Oversubscription is a common practice for improving cloud resource utilization. It allows the cloud service provider to sell more resources than the physical limit, assuming not all users would fully utilize the resources simultaneously. However, how to design an oversubscription policy that improves utilization while satisfying the some safety constraints remains an open problem. Existing methods and industrial practices are over-conservative, ignoring the coordination of diverse resource usage patterns and probabilistic constraints. To address these two limitations, this paper formulates the oversubscription for cloud as a chance-constrained optimization problem and propose an effective Chance-Constrained Multi-Agent Reinforcement Learning (C2MARL) method to solve this problem. Specifically, C2MARL reduces the number of constraints by considering their upper bounds and leverages a multi-agent reinforcement learning paradigm to learn a safe and optimal coordination policy. We evaluate our C2MARL on an internal cloud platform and public cloud datasets. Experiments show that our C2MARL outperforms existing methods in improving utilization (20% ∼ 86%) under different levels of safety constraints.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Coach: Exploiting Temporal Patterns for All-Resource Oversubscription in Cloud PlatformsBenjamin Reidys, Pantea Zardoshti, Íñigo Goiri, Celine Irvene 等ASPLOS 2025 · 被引用 9 次
- On the Hardness of Constrained Cooperative Multi-Agent Reinforcement LearningZiyi Chen, Yi Zhou, Heng HuangICLR 2024 · 被引用 6 次
它引用的顶会 Paper3
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud ProviderMohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Irfan Chaudhry 等USENIX ATC 2020 · 被引用 946 次
- Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision ProcessesDongsheng Ding, Kaiqing Zhang, Tamer Basar, Mihailo R. JovanovicNeurIPS 2020 · 被引用 252 次
- Protean: VM Allocation Service at ScaleOri Hadary, Luke Marshall, Ishai Menache, Abhisek Pan 等OSDI 2020 · 被引用 189 次
相关 Paper
- AWARE: Automate Workload Autoscaling with Reinforcement Learning in Production Cloud SystemsHaoran Qiu, Weichao Mao, Chen Wang, Hubertus Franke 等USENIX ATC 2023 · 被引用 95 次
- Following the Usage, Not the Request: Risk-Aware Task Scheduling with Overbooking in Edge CloudsTie Ma, Shan Zhang, Xiaoyu Zhang, Zichuan Zheng 等INFOCOM 2026
- Take it to the limit: peak prediction-driven resource overcommitment in datacentersNoman Bashir, Nan Deng, Krzysztof Rzadca, David Irwin 等EuroSys 2021 · 被引用 60 次
- Flipping-based Policy for Chance-Constrained Markov Decision ProcessesXun Shen, Shuo Jiang, Akifumi Wachi, Kazumune Hashimoto 等NeurIPS 2024 · 被引用 5 次
- Multi-Unit Auctions for Allocating Chance-Constrained ResourcesAnna Gautier, Bruno Lacerda, Nick Hawes, Michael J. WooldridgeAAAI 2023 · 被引用 5 次
