Cilantro: Performance-Aware Resource Allocation for General Objectives via Online Feedback
Romil Bhardwaj, Kirthevasan Kandasamy, Asim Biswal, Wenshuo Guo, Benjamin Hindman, Joseph Gonzalez, Michael I. Jordan, Ion Stoica
Abstract
Traditional systems for allocating finite cluster resources among competing jobs have either aimed at providing fairness, relied on users to specify their resource requirements, or have estimated these requirements via surrogate metrics (e.g. CPU utilization). These approaches do not account for a job's real world performance (e.g. P95 latency). Existing performance-aware systems use offline profiled data and/or are designed for specific allocation objectives. In this work, we argue that resource allocation systems should directly account for real-world performance and the varied allocation objectives of users. In this pursuit, we build Cilantro.
At the core of Cilantro is an online learning mechanism which forms feedback loops with the jobs to estimate the resource to performance mappings and load shifts. This relieves users from the onerous task of job profiling and collects reliable real-time feedback. This is then used to achieve a variety of user-specified scheduling objectives. Cilantro handles the uncertainty in the learned models by adapting the underlying policy to work with confidence bounds. We demonstrate this in two settings. First, in a multi-tenant 1000 CPU cluster with 20 independent jobs, three of Cilantro's policies outperform 9 other baselines on three different performance-aware scheduling objectives, improving user utilities by up to 1.2 -3.7× and performs comparably to oracular policies. Second, in a microservices setting, where 160 CPUs must be distributed between 19 inter-dependent microservices, Cilantro outperforms 3 other baselines, reducing the end-to-end P99 latency to ×0.57 the next best baseline.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b52fd6b-fa58-4407-a493-bd08bdb75b97Cited by top-tier papers9
- Harmonizing Efficiency and Practicability: Optimizing Resource Utilization in Serverless Computing with JiaguQingyuan Liu, Yanning Yang, Dong Du, Yubin Xia et al.USENIX ATC 2024 · 39 citations
- Jolteon: Unleashing the Promise of Serverless for Serverless WorkflowsZili Zhang, Chao Jin, Xin JinNSDI 2024 · 15 citations
- TopFull: An Adaptive Top-Down Overload Control for SLO-Oriented MicroservicesJinwoo Park, Jaehyeong Park, Youngmok Jung, Hwijoon Lim et al.SIGCOMM 2024 · 10 citations
- CAPSys: Contention-aware task placement for data stream processingYuanli Wang, Lei Huang, Zikun Wang, Vasiliki Kalavri et al.EuroSys 2025 · 8 citations
- When will my ML Job finish? Toward providing Completion Time Estimates through Predictability-Centric SchedulingAbdullah Bin Faisal, Noah Martin, Hafiz Mohsin Bashir, Swaminathan Lamelas et al.OSDI 2024 · 6 citations
Builds on5
- FIRM: An Intelligent Fine-grained Resource Management Framework for SLO-Oriented MicroservicesHaoran Qiu, Subho S. Banerjee, Saurabh Jha, Zbigniew T. Kalbarczyk et al.OSDI 2020 · 350 citations
- Autopilot: workload autoscaling at GoogleKrzysztof Rzadca, Pawel Findeisen, Jacek Swiderski, Przemyslaw Zych et al.EuroSys 2020 · 299 citations
- Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning WorkloadsDeepak Narayanan, Keshav Santhanam, Fiodar Kazhamiaka, Amar Phanishayee et al.OSDI 2020 · 286 citations
- Sinan: ML-based and QoS-aware resource management for cloud microservicesYanqi Zhang, Weizhe Hua, Zhuangzhuang Zhou, G. Edward Suh et al.ASPLOS 2021 · 226 citations
- RackSched: A Microsecond-Scale Scheduler for Rack-Scale ComputersHang Zhu, Kostis Kaffes, Zixu Chen, Zhenming Liu et al.OSDI 2020 · 58 citations
Related papers
- A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via FaroBeomyeol Jeon, Chen Wang, Diana Arroyo, Alaa Youssef et al.EuroSys 2025
- Spark-based Cloud Data Analytics using Multi-Objective OptimizationFei Song, Khaled Zaouk, Chenghao Lyu, Arnab Sinha et al.ICDE 2021 · 15 citations
- CLITE: Efficient and QoS-Aware Co-Location of Multiple Latency-Critical Jobs for Warehouse Scale ComputersTirthak Patel, Devesh TiwariHPCA 2020 · 153 citations
- Harvesting Spare CPU Resources in Container SystemsAdam Hall, Anirudh Sarma, Esha Choukse, Umakishore Ramachandran et al.NSDI 2026 · 2 citations
- Autothrottle: A Practical Bi-Level Approach to Resource Management for SLO-Targeted MicroservicesZibo Wang, Pinghe Li, Chieh-Jan Mike Liang, Feng Wu et al.NSDI 2024
