Kairos: Building Cost-Efficient Machine Learning Inference Systems with Heterogeneous Cloud Resources
Baolin Li, Siddharth Samsi, Vijay Gadepally, Devesh Tiwari
摘要
Online inference is becoming a key service product for many businesses, deployed in cloud platforms to meet customer demands. Despite their revenue-generation capability, these services need to operate under tight Quality-of-Service (QoS) and cost budget constraints. This paper introduces Kairos 1 , a novel runtime framework that maximizes the query throughput while meeting QoS target and a cost budget. Kairos designs and implements novel techniques to build a pool of heterogeneous compute hardware without online exploration overhead, and distribute inference queries optimally at runtime. Our evaluation using industry-grade machine learning (ML) models shows that Kairos yields up to 2× the throughput of an optimal homogeneous solution, and outperforms state-of-the-art schemes by up to 70%, despite advantageous implementations of the competing schemes to ignore their exploration overhead.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Clover: Toward Sustainable AI with Carbon-Aware Machine Learning Inference ServiceBaolin Li, Siddharth Samsi, Vijay Gadepally, Devesh TiwariSC 2023 · 被引用 63 次
- Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy ScalingSohaib Ahmad, Hui Guan, Ramesh K. SitaramanHPDC 2024 · 被引用 8 次
- SMIless: Serving DAG-based Inference with Dynamic Invocations under Serverless ComputingChengzhi Lu, Huanle Xu, Yudan Li, Wenyan Chen 等SC 2024 · 被引用 8 次
它引用的顶会 Paper17
- MLPerf Inference BenchmarkVijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson 等ISCA 2020 · 被引用 517 次
- Serving DNNs like Clockwork: Performance Predictability from the Bottom UpArpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao 等OSDI 2020 · 被引用 392 次
- Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning WorkloadsDeepak Narayanan, Keshav Santhanam, Fiodar Kazhamiaka, Amar Phanishayee 等OSDI 2020 · 被引用 286 次
- Sinan: ML-based and QoS-aware resource management for cloud microservicesYanqi Zhang, Weizhe Hua, Zhuangzhuang Zhou, G. Edward Suh 等ASPLOS 2021 · 被引用 226 次
- Batch: machine learning inference serving on serverless platforms with adaptive batchingAhsan Ali, Riccardo Pinciroli, Feng Yan, Evgenia SmirniSC 2020 · 被引用 184 次
相关 Paper
- Compile-Time QoS Scheme for Deep Learning InferencesSungin Hong, Hyunjun Kim, Hwansoo HanSC 2025 · 被引用 1 次
- RIBBON: cost-effective and qos-aware deep learning model inference using a diverse pool of cloud computing instancesBaolin Li, Rohan Basu Roy, Tirthak Patel, Vijay Gadepally 等SC 2021 · 被引用 16 次
- Kalmia: A Heterogeneous QoS-aware Scheduling Framework for DNN Tasks on Edge ServersZiyan Fu, Ju Ren, Deyu Zhang, Yuezhi Zhou 等INFOCOM 2022 · 被引用 28 次
- A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via FaroBeomyeol Jeon, Chen Wang, Diana Arroyo, Alaa Youssef 等EuroSys 2025
- Proteus: A High-Throughput Inference-Serving System with Accuracy ScalingSohaib Ahmad, Hui Guan, Brian D. Friedman, Thomas Williams 等ASPLOS 2024 · 被引用 31 次
