SOL: safe on-node learning in cloud platforms
Yawen Wang, Daniel Crankshaw, Neeraja J. Yadwadkar, Daniel S. Berger, Christos Kozyrakis, Ricardo Bianchini
Abstract
Cloud platforms run many software agents on each server node. These agents manage all aspects of node operation, and in some cases frequently collect data and make decisions. Unfortunately, their behavior is typically based on pre-defined static heuristics or offline analysis; they do not leverage on-node machine learning (ML). In this paper, we first characterize the spectrum of node agents in Azure, and identify the classes of agents that are most likely to benefit from on-node ML. We then propose SOL, an extensible framework for designing ML-based agents that are safe and robust to the range of failure conditions that occur in production. SOL provides a simple API to agent developers and manages the scheduling and running of the agent-specific functions they write. We illustrate the use of SOL by implementing three ML-based agents that manage CPU cores, node power, and memory placement. Our experiments show that (1) ML substantially improves our agents, and (2) SOL ensures that agents operate safely under a variety of failure conditions. We conclude that ML-based agents show significant potential and that SOL can help build them.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4a250200-1514-4dd1-bda5-efd8e30b3f4dCited by top-tier papers6
- AWARE: Automate Workload Autoscaling with Reinforcement Learning in Production Cloud SystemsHaoran Qiu, Weichao Mao, Chen Wang, Hubertus Franke et al.USENIX ATC 2023 · 95 citations
- Power-aware Deep Learning Model Serving with μ-ServeHaoran Qiu, Weichao Mao, Archit Patke, Shengkun Cui et al.USENIX ATC 2024 · 82 citations
- Towards a Machine Learning-Assisted Kernel with LAKEHenrique Fingler, Isha Tarte, Hangchen Yu, Ariel Szekely et al.ASPLOS 2023 · 19 citations
- Wave: Offloading Resource Management to SmartNIC CoresJack Tigar Humphries, Neel Natu, Kostis Kaffes, Stanko Novakovic et al.ASPLOS 2025 · 5 citations
- FleetIO: Managing Multi-Tenant Cloud Storage with Multi-Agent Reinforcement LearningJinghan Sun, Benjamin Reidys, Daixuan Li, Jichuan Chang et al.ASPLOS 2025 · 5 citations
Builds on8
- Protean: VM Allocation Service at ScaleOri Hadary, Luke Marshall, Ishai Menache, Abhisek Pan et al.OSDI 2020 · 189 citations
- Exploring the Design Space of Page Management for Multi-Tiered Memory SystemsJonghyeon Kim, Wonkyo Choe, Jeongseob AhnUSENIX ATC 2021 · 108 citations
- LinnOS: Predictability on Unpredictable Flash Storage with a Light Neural NetworkMingzhe Hao, Levent Toksoz, Nanqinqin Li, Edward Edberg Halim et al.OSDI 2020 · 97 citations
- Prediction-Based Power Oversubscription in Cloud PlatformsAlok Gautam Kumbhare, Reza Azimi, Ioannis Manousakis, Anand Bonde et al.USENIX ATC 2021 · 90 citations
- Adaptive low-overhead scheduling for periodic and reactive intermittent executionKiwan Maeng, Brandon LuciaPLDI 2020 · 84 citations
Related papers
- A Dual-Agent Scheduler for Distributed Deep Learning Jobs on Public Cloud via Reinforcement LearningMingzhe Xing, Hangyu Mao, Shenglin Yin, Lichen Pan et al.KDD 2023 · 9 citations
- AutoSys: The Design and Operation of Learning-Augmented SystemsChieh-Jan Mike Liang, Hui Xue, Mao Yang, Lidong Zhou et al.USENIX ATC 2020 · 17 citations
- Predictive and Adaptive Failure Mitigation to Avert Production Cloud VM InterruptionsSebastien Levy, Randolph Yao, Youjiang Wu, Yingnong Dang et al.OSDI 2020 · 35 citations
- Memory-harvesting VMs in cloud platformsAlexander Fuerst, Stanko Novakovic, Iñigo Goiri, Gohar Irfan Chaudhry et al.ASPLOS 2022 · 39 citations
- Runtime Variation in Big Data AnalyticsYiwen Zhu, Rathijit Sen, Robert Horton, John Mark AgostaSIGMOD 2023 · 5 citations
