A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via Faro
Beomyeol Jeon, Chen Wang, Diana Arroyo, Alaa Youssef, Indranil Gupta
摘要
This paper tackles the challenge of running multiple ML inference jobs (models) under time-varying workloads, on a constrained on-premises production cluster. Our system Faro takes in latency Service Level Objectives (SLOs) for each job, auto-distills them into utility functions, "sloppifies" these utility functions to make them amenable to mathematical optimization, automatically predicts workload via probabilistic prediction, and dynamically makes implicit cross-job resource allocations, in order to satisfy cluster-wide objectives, e.g., total utility, fairness, and other hybrid variants. A major challenge Faro tackles is that using precise utilities and high-fidelity predictors, can be too slow (and in a sense too precise!) for the fast adaptation we require. Faro's solution is to "sloppify" (relax) its multiple design components to achieve fast adaptation without overly degrading solution quality. Faro is implemented in a stack consisting of Ray Serve running atop a Kubernetes cluster. Trace-driven cluster deployments show that Faro achieves 2.3×-23× lower SLO violations compared to state-of-the-art systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud ProviderMohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Irfan Chaudhry 等USENIX ATC 2020 · 被引用 946 次
- MLPerf Inference BenchmarkVijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson 等ISCA 2020 · 被引用 517 次
- NHITS: Neural Hierarchical Interpolation for Time Series ForecastingCristian Challu, Kin G. Olivares, Boris N. Oreshkin, Federico Garza Ramírez 等AAAI 2023 · 被引用 420 次
- Serving DNNs like Clockwork: Performance Predictability from the Bottom UpArpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao 等OSDI 2020 · 被引用 392 次
- INFaaS: Automated Model-less Inference ServingFrancisco Romero, Qian Li, Neeraja J. Yadwadkar, Christos KozyrakisUSENIX ATC 2021 · 被引用 325 次
相关 Paper
- QoServe: Breaking the Silos of LLM Inference ServingKanishk Goel, Jayashree Mohan, Nipun Kwatra, Ravi Shreyas Anupindi 等ASPLOS 2026 · 被引用 3 次
- NeuRO: Inference-time Profiling and Orchestration of ML Applications at the EdgeArshad Javeed, György Dán, Viktoria FodorINFOCOM 2026 · 被引用 1 次
- Multiplexing Dynamic Deep Learning Workloads with SLO-awareness in GPU ClustersWenyan Chen, Chengzhi Lu, Huanle Xu, Kejiang Ye 等EuroSys 2025 · 被引用 13 次
- HyGen: Efficient LLM Serving via Elastic Online-Offline Request Co-locationTing Sun, Penghan Wang, Fan LaiNeurIPS 2025 · 被引用 17 次
- Cilantro: Performance-Aware Resource Allocation for General Objectives via Online FeedbackRomil Bhardwaj, Kirthevasan Kandasamy, Asim Biswal, Wenshuo Guo 等OSDI 2023 · 被引用 41 次
