SelfTune: Tuning Cluster Managers
Ajaykrishna Karthikeyan, Nagarajan Natarajan, Gagan Somashekar, Lei Zhao, Ranjita Bhagwan, Rodrigo Fonseca, Tatiana Racheva, Yogesh Bansal
Abstract
Large-scale cloud providers rely on cluster managers for container allocation and load balancing (e.g., Kubernetes), VM provisioning (e.g., Protean), and other management tasks. These cluster managers use algorithms or heuristics whose behavior depends upon multiple configuration parameters. Currently, operators manually set these parameters using a combination of domain knowledge and limited testing. In very large-scale and dynamic environments, these manually-set parameters may lead to sub-optimal cluster states, adversely affecting important metrics such as latency and throughput.
In this paper we describe SelfTune, a framework that automatically tunes such parameters in deployment. SelfTune piggybacks on the iterative nature of cluster managers which, through multiple iterations, drives a cluster to a desired state. Using a simple interface, developers integrate SelfTune into the cluster manager code, which then uses a principled reinforcement learning algorithm to tune important parameters over time. We have deployed SelfTune on tens of thousands of machines that run a large-scale background task scheduler at Microsoft. SelfTune has improved throughput by as much as 20% in this deployment by continuously tuning a key configuration parameter that determines the number of jobs concurrently accessing CPU and disk on every machine. We also evaluate SelfTune with two Azure FaaS workloads, the Kubernetes Vertical Pod Autoscaler, and the DeathStar microservice benchmark. In all cases, SelfTune significantly improves cluster performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4cd4452d-f6a2-46ce-ae7a-8e90d899171bCited by top-tier papers7
- OPPerTune: Post-Deployment Configuration Tuning of Services Made EasyGagan Somashekar, Karan Tandon, Anush Kini, Chieh-Chun Chang et al.NSDI 2024 · 24 citations
- When will my ML Job finish? Toward providing Completion Time Estimates through Predictability-Centric SchedulingAbdullah Bin Faisal, Noah Martin, Hafiz Mohsin Bashir, Swaminathan Lamelas et al.OSDI 2024 · 6 citations
- TraceUpscaler: Upscaling Traces to Evaluate Systems at High LoadSultan Mahmud Sajal, Timothy Zhu, Bhuvan Urgaonkar, Siddhartha SenEuroSys 2024 · 4 citations
- The Same Only Different: On Information Modality for Configuration Performance AnalysisHongyuan Liang, Yue Huang, Tao ChenICSE 2025 · 3 citations
- Kamino: Efficient VM Allocation at Scale with Latency-Driven Cache-Aware SchedulingDavid Domingo, Hugo Barbalho, Marco Molinaro, Kuan Liu et al.OSDI 2025 · 2 citations
Builds on9
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud ProviderMohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Irfan Chaudhry et al.USENIX ATC 2020 · 946 citations
- FIRM: An Intelligent Fine-grained Resource Management Framework for SLO-Oriented MicroservicesHaoran Qiu, Subho S. Banerjee, Saurabh Jha, Zbigniew T. Kalbarczyk et al.OSDI 2020 · 350 citations
- Autopilot: workload autoscaling at GoogleKrzysztof Rzadca, Pawel Findeisen, Jacek Swiderski, Przemyslaw Zych et al.EuroSys 2020 · 299 citations
- Protean: VM Allocation Service at ScaleOri Hadary, Luke Marshall, Ishai Menache, Abhisek Pan et al.OSDI 2020 · 189 citations
- Twine: A Unified Cluster Management System for Shared InfrastructureChunqiang Tang, Kenny Yu, Kaushik Veeraraghavan, Jonathan Kaldor et al.OSDI 2020 · 107 citations
Related papers
- Metis: learning to schedule long-running applications in shared container clusters at scaleLuping Wang, Qizhen Weng, Wei Wang, Chen Chen et al.SC 2020 · 46 citations
- EdgeTuner: Fast Scheduling Algorithm Tuning for Dynamic Edge-Cloud Workloads and ResourcesRui Han, Shilin Wen, Chi Harold Liu, Ye Yuan et al.INFOCOM 2022 · 27 citations
- ResTune: Resource Oriented Tuning Boosted by Meta-Learning for Cloud DatabasesXinyi Zhang, Hong Wu, Zhuo Chang, Shuowei Jin et al.SIGMOD 2021 · 113 citations
- Scarf: Self-Adaptive Tuning via Multi-Objective Reinforcement Learning for Apache FlinkLiu Liu, Shenghao Gong, Ziquan Fang, Yunjun GaoVLDB 2026
- TuneAgent: Agentic Operating System Kernel Tuning with Reinforcement LearningHongyu Lin, Yuchen Li, Haoran Luo, Zhenghong Lin et al.KDD 2026 · 3 citations
