Intelligent Resource Scheduling for Co-located Latency-critical Services: A Multi-Model Collaborative Learning Approach
Lei Liu, Xinglei Dou, Yuetao Chen
Abstract
Latency-critical services have been widely deployed in cloud environments. For cost-efficiency, multiple services are usually co-located on a server. Thus, run-time resource scheduling becomes the pivot for QoS control in these complicated co-location cases. However, the scheduling exploration space enlarges rapidly with the increasing server resources, making the schedulers hardly provide ideal solutions quickly. More importantly, we observe that there are "resource cliffs" in the scheduling exploration space. They affect the exploration efficiency and always lead to severe QoS fluctuations. Resource cliffs cannot be easily avoided in previous schedulers. To address these problems, we propose a novel ML-based intelligent scheduler -OSML. It learns the correlation between architectural hints (e.g., IPC, cache misses, memory footprint, etc.), scheduling solutions and the QoS demands based on a data set we collected from 11 widely deployed services running on off-the-shelf servers. OSML employs multiple ML models to work collaboratively to predict QoS variations, shepherd the scheduling, and recover from QoS violations in complicated co-location cases. OSML can intelligently avoid resource cliffs during scheduling and reach an optimal solution much faster than previous approaches for co-located LC services. Experimental results show that OSML supports higher loads and meets QoS targets with lower scheduling overheads and shorter convergence time than previous studies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Enabling SLO-Aware 5G Multi-Access Edge Computing with SMECXiao Zhang, Daehyeok KimNSDI 2026 · 4 citations
- Heimdall: Optimizing Storage I/O Admission with Extensive Machine Learning PipelineDaniar Heri Kurniawan, Rani Ayu Putri, Peiran Qin, Kahfi S. Zulkifli et al.EuroSys 2025 · 3 citations
- MedFS: Pursuing Low Update Overhead via Metadata-Enabled Delta Compression for Log-structured File System on Mobile DeviceChao Wu, Cheng Ji, Li-Pin Chang, Zongwei Zhu et al.FAST 2025
Builds on5
- Sinan: ML-based and QoS-aware resource management for cloud microservicesYanqi Zhang, Weizhe Hua, Zhuangzhuang Zhou, G. Edward Suh et al.ASPLOS 2021 · 226 citations
- Caladan: Mitigating Interference at Microsecond TimescalesJoshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam BelayOSDI 2020 · 213 citations
- CLITE: Efficient and QoS-Aware Co-Location of Multiple Latency-Critical Jobs for Warehouse Scale ComputersTirthak Patel, Devesh TiwariHPCA 2020 · 153 citations
- Twig: Multi-Agent Task Management for Colocated Latency-Critical Cloud ServicesRajiv Nishtala, Vinicius Petrucci, Paul M. Carpenter, Magnus SjälanderHPCA 2020 · 76 citations
- CuttleSys: Data-Driven Resource Management for Interactive Services on Reconfigurable MulticoresNeeraj Kulkarni, Gonzalo Gonzalez-Pumariega, Amulya Khurana, Christine A. Shoemaker et al.MICRO 2020 · 21 citations
Related papers
- Dynamic Edge-centric Resource Provisioning for Online and Offline Services Co-locationTao Ouyang, Kongyange Zhao, Xiaoxi Zhang, Zhi Zhou et al.INFOCOM 2023 · 18 citations
- OLPart: Online Learning based Resource Partitioning for Colocating Multiple Latency-Critical Jobs on Commodity ComputersRuobing Chen, Haosen Shi, Yusen Li, Xiaoguang Liu et al.EuroSys 2023 · 27 citations
- NeuRO: Inference-time Profiling and Orchestration of ML Applications at the EdgeArshad Javeed, György Dán, Viktoria FodorINFOCOM 2026 · 1 citation
- UFO: The Ultimate QoS-Aware Core Management for Virtualized and Oversubscribed Public CloudsYajuan Peng, Shuang Chen, Yi Zhao, Zhibin YuNSDI 2024 · 7 citations
- Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal SharingSeungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park et al.USENIX ATC 2022 · 200 citations
