Derm: SLA-aware Resource Management for Highly Dynamic Microservices
Liao Chen, Shutian Luo, Chenyu Lin, Zizhao Mo, Huanle Xu, Kejiang Ye, ChengZhong Xu
摘要
Ensuring efficient resource allocation while providing service level agreement (SLA) guarantees for end-to-end (E2E) latency is crucial for microservice applications. Although existing studies have made significant contributions towards achieving this objective, they primarily concentrate on static graphs. However, microservice graphs are inherently dynamic during runtime in production environments, necessitating more effective and scalable resource management solutions.In this paper, we present Derm, a new resource management system designed for microservice applications with highly dynamic graphs. Our principal finding is that prioritizing different microservice graphs can lead to a substantial reduction in resource allocation. To take advantage of this opportunity, we develop three main components. The first is a performance model that describes uncertainties of microservice latency through a conditional exponential distribution. The second is a probabilistic quantification of the dynamics of microservice graphs. The third is an optimization method for adjusting the resource allocation of microservices to minimize resource usage. We evaluate Derm in our cluster using real microservice benchmarks and production traces. The results highlight that Derm reduces the resource usage by and lowers SLA violation probability by , compared to existing approaches.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Fast and Efficient Scaling for Microservices with SurgeGuardAnyesha Ghosh, Neeraja J. Yadwadkar, Mattan ErezSC 2024 · 被引用 3 次
- Ursa: Lightweight Resource Management for Cloud-Native MicroservicesYanqi Zhang, Zhuangzhuang Zhou, Sameh Elnikety, Christina DelimitrouHPCA 2024 · 被引用 14 次
- Grad: Intelligent Microservice Scaling by Harnessing Resource FungibilityLiao Chen, Chenyu Lin, Shutian Luo, Huanle Xu 等HPCA 2025 · 被引用 5 次
- Nodens: Enabling Resource Efficient and Fast QoS Recovery of Dynamic Microservice Applications in DatacentersJiuchen Shi, Hang Zhang, Zhixin Tong, Quan Chen 等USENIX ATC 2023 · 被引用 32 次
- PERT-GNN: Latency Prediction for Microservice-based Cloud-Native Applications via Graph Neural NetworksDa Sun Handason Tam, Yang Liu, Huanle Xu, Siyue Xie 等KDD 2023 · 被引用 20 次
