AWARE: Automate Workload Autoscaling with Reinforcement Learning in Production Cloud Systems
Haoran Qiu, Weichao Mao, Chen Wang, Hubertus Franke, Alaa Youssef, Zbigniew T. Kalbarczyk, Tamer Basar, Ravishankar K. Iyer
Abstract
Workload autoscaling is widely used in public and private cloud systems to maintain stable service performance and save resources. However, it remains challenging to set the optimal resource limits and dynamically scale each workload at runtime. Reinforcement learning (RL) has recently been proposed and applied in various systems tasks, including resource management. In this paper, we first characterize the state-of-the-art RL approaches for workload autoscaling in a public cloud and point out that there is still a large gap in taking the RL advances to production systems. We then propose AWARE, an extensible framework for deploying and managing RL-based agents in production systems. AWARE leverages meta-learning and bootstrapping to (a) automatically and quickly adapt to different workloads, and (b) provide safe and robust RL exploration. AWARE provides a common OpenAI Gym-like RL interface to agent developers for easy integration with different systems tasks. We illustrate the use of AWARE in the case of workload autoscaling. Our experiments show that AWARE adapts a learned autoscaling policy to new workloads 5.5× faster than the existing transfer-learning-based approach and provides stable online policy-serving performance with less than 3.6% reward degradation. With bootstrapping, AWARE helps achieve 47.5% and 39.2% higher CPU and memory utilization while reducing SLO violations by a factor of 16.9× during policy training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 33d45909-28df-4585-b9c4-0709040c8a25Cited by top-tier papers11
- Power-aware Deep Learning Model Serving with μ-ServeHaoran Qiu, Weichao Mao, Archit Patke, Shengkun Cui et al.USENIX ATC 2024 · 82 citations
- Designing Cloud Servers for Lower CarbonJaylen Wang, Daniel S. Berger, Fiodar Kazhamiaka, Celine Irvene et al.ISCA 2024 · 49 citations
- Harmonizing Efficiency and Practicability: Optimizing Resource Utilization in Serverless Computing with JiaguQingyuan Liu, Yanning Yang, Dong Du, Yubin Xia et al.USENIX ATC 2024 · 39 citations
- FLOAT: Federated Learning Optimizations with Automated TuningAhmad Faraz Khan, Azal Ahmad Khan, Ahmed M. Abdelmoniem, Samuel Fountain et al.EuroSys 2024 · 22 citations
- Multi-Agent Meta-Reinforcement Learning: Sharper Convergence Rates with Task SimilarityWeichao Mao, Haoran Qiu, Chen Wang, Hubertus Franke et al.NeurIPS 2023 · 17 citations
Builds on9
- Learning in situ: a randomized experiment in video streamingFrancis Y. Yan, Hudson Ayers, Chenzhi Zhu, Sadjad Fouladi et al.NSDI 2020 · 360 citations
- FIRM: An Intelligent Fine-grained Resource Management Framework for SLO-Oriented MicroservicesHaoran Qiu, Subho S. Banerjee, Saurabh Jha, Zbigniew T. Kalbarczyk et al.OSDI 2020 · 350 citations
- Autopilot: workload autoscaling at GoogleKrzysztof Rzadca, Pawel Findeisen, Jacek Swiderski, Przemyslaw Zych et al.EuroSys 2020 · 299 citations
- Faster and Cheaper Serverless Computing on Harvested ResourcesYanqi Zhang, Iñigo Goiri, Gohar Irfan Chaudhry, Rodrigo Fonseca et al.SOSP 2021 · 131 citations
- Accelerometer: Understanding Acceleration Opportunities for Data Center Overheads at HyperscaleAkshitha Sriraman, Abhishek DhanotiaASPLOS 2020 · 78 citations
Related papers
- Metis: learning to schedule long-running applications in shared container clusters at scaleLuping Wang, Qizhen Weng, Wei Wang, Chen Chen et al.SC 2020 · 46 citations
- A Dual-Agent Scheduler for Distributed Deep Learning Jobs on Public Cloud via Reinforcement LearningMingzhe Xing, Hangyu Mao, Shenglin Yin, Lichen Pan et al.KDD 2023 · 9 citations
- Graph Assisted Offline-Online Deep Reinforcement Learning for Dynamic Workflow SchedulingYifan Yang, Gang Chen, Hui Ma, Cong Zhang et al.ICLR 2025
- RLBoost: Harvesting Preemptible Cloud Resources for Cost-Efficient Reinforcement Learning on LLMsYongji Wu, Xueshen Liu, Haizhong Zheng, Juncheng Gu et al.NSDI 2026 · 4 citations
- Practical Efficient Microservice Autoscaling with QoS AssuranceMd Rajib Hossen, Mohammad A. Islam, Kishwar AhmedHPDC 2022 · 45 citations
