Erlang: Application-Aware Autoscaling for Cloud Microservices
Vighnesh Sachidananda, Anirudh Sivaraman
摘要
As cloud applications shift from monoliths to loosely coupled microservices, application developers must decide how many compute resources (e.g., number of replicated containers) to assign to each microservice within an application. This decision affects both (1) the dollar cost to the application developer and (2) the end-to-end latency perceived by the application user. Today, individual microservices are autoscaled independently by adding VMs whenever per-microservice CPU or memory utilization crosses a configurable threshold. However, an application user's end-to-end latency consists of time spent on multiple microservices and each microservice might need a different number of VMs to achieve an overall end-to-end latency.
We present Erlang, an autoscaler for microservice-based applications, which collectively allocates VMs to microservices with a global goal of minimizing dollar cost while keeping end-to-end application latency under a given target. Using 5 open-source applications, we compared Erlang to several utilization and machine learning based autoscalers. We evaluate Erlang across different compute settings on Google Kubernetes Engine (GKE) in which users manage compute resources, GKE standard, and a new mode of operation in which the cloud provider manages compute infrastructure, GKE Autopilot. Erlang meets a desired median or tail latency target on 53 of 63 workloads where it provides a cost reduction of 19.3%, on average, over the next cheapest autoscaler. Erlang is the most cost effective autoscaling policy for 48 of these 53 workloads. The cost savings from managing a cluster with Erlang result in Erlang paying for its training cost in a few days. On smaller applications, *Work done while at Stanford University.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- JITServe: SLO-aware LLM Serving with Imprecise Request InformationWei Zhang, Zhiyu Wu, Yi Mu, Rui Ning 等NSDI 2026 · 被引用 29 次
- Towards Performance Robustness for MicroservicesDivyanshu Saxena, Gaurav Vipat, Jiaxin Lin, Jingbo Wang 等NSDI 2026
- MerKury: Adaptive Resource Allocation to Enhance the Kubernetes Performance for Large-Scale ClustersJiayin Luo, Xinkui Zhao, Yuxin Ma, Shengye Pang 等WWW 2025
- CONGO: Compressive Online Gradient OptimizationJeremy Carleton, Prathik Vijaykumar, Divyanshu Saxena, Dheeraj Narasimha 等ICLR 2025
它引用的顶会 Paper3
- FIRM: An Intelligent Fine-grained Resource Management Framework for SLO-Oriented MicroservicesHaoran Qiu, Subho S. Banerjee, Saurabh Jha, Zbigniew T. Kalbarczyk 等OSDI 2020 · 被引用 350 次
- Autopilot: workload autoscaling at GoogleKrzysztof Rzadca, Pawel Findeisen, Jacek Swiderski, Przemyslaw Zych 等EuroSys 2020 · 被引用 299 次
- Sinan: ML-based and QoS-aware resource management for cloud microservicesYanqi Zhang, Weizhe Hua, Zhuangzhuang Zhou, G. Edward Suh 等ASPLOS 2021 · 被引用 226 次
相关 Paper
- Autothrottle: A Practical Bi-Level Approach to Resource Management for SLO-Targeted MicroservicesZibo Wang, Pinghe Li, Chieh-Jan Mike Liang, Feng Wu 等NSDI 2024
- Practical Efficient Microservice Autoscaling with QoS AssuranceMd Rajib Hossen, Mohammad A. Islam, Kishwar AhmedHPDC 2022 · 被引用 45 次
- DeepScaler: Holistic Autoscaling for Microservices Based on Spatiotemporal GNN with Adaptive Graph LearningChunyang Meng, Shijie Song, Haogang Tong, Maolin Pan 等ASE 2023 · 被引用 28 次
- Astraea: towards QoS-aware and resource-efficient multi-stage GPU servicesWei Zhang, Quan Chen, Kaihua Fu, Ningxin Zheng 等ASPLOS 2022 · 被引用 28 次
- SLATE: Service Layer Traffic EngineeringGangmuk Lim, Aditya Prerepa, Brighten Godfrey, Radhika MittalNSDI 2026
