ReTail: Opting for Learning Simplicity to Enable QoS-Aware Power Management in the Cloud
Shuang Chen, Angela Jin, Christina Delimitrou, José F. Martínez
摘要
Many cloud services have Quality-of-Service (QoS) requirements; most requests have to to complete within a given latency constraint. Recently, researchers have begun to investigate whether it is possible to meet QoS while attempting to save power on a per-request basis. Existing work shows that one can indeed hand-tune a request latency predictor offline for a particular cloud application, and consult it at runtime to modulate CPU voltage and frequency, resulting in substantial power savings.
In this paper, we propose ReTail, an automated and general solution for request-level power management of latency-critical services with QoS constraints. We present a systematic process to select the features of any given application that best correlate with its request latency. ReTail uses these features to predict latency, and adjust CPU's power consumption. ReTail's predictor is trained fully at runtime. We show that unlike previous findings, simple techniques perform better than complex machine learning models, when using the right input features. For a web search engine, ReTail outperforms prior mechanisms based on complex hand-tuned predictors for that application domain. Furthermore, ReTail's systematic approach also yields superior power savings across a diverse set of cloud applications. TABLE I: Qualitative comparison of ReTail vs. two closely related proposals for power management of latency-critical applications. Rubik Gemini ReTail Method Statistical model Neural net Linear regression Feature space N/A Request Request&Application Feature selection N/A Hand-picked Systematic Request-accurate General applications Training overhead Low High Low Inference overhead Low High Low QoS guarantee No dropped requests Adapt to model drift dynamically adjust to system or application interference by retraining the latency predictor online. II. RELATED WORK Most proposals towards improving power efficiency [15, 26, 43] focus on throughput-oriented batch jobs. Recent work that explores DVFS for LC services with latency/QoS constraints can be broadly classified into two categories: Coarse-grained power management like Pegasus [34] dynamically adjusts frequency for the entire LC application. It addresses long-term latency variations under load fluctuations, but leaves power savings on the table by not differentiating individual requests.
Fine-grained power management outperforms coarsegrained management by differentiating at request granularity, generally boosting long and slowing down short requests, while trying to meet QoS. There are two main methods:
• Classification-based methods classify requests into short and long using various metrics, and boost all requests in the long category. EETL [23] tracks the progress of each request; those that exceed a predetermined execution threshold are flagged as long requests, and EETL boosts the frequency at that point. However, by the time a request reaches the progress threshold, it may be too late to prevent tail latency degradation. Adrenaline [24] introduces the concept of featuredriven request classification. It studies two LC applications, web search and key-value stores, and identifies, by human inspection, request features in each that can be used to predict incoming requests as being either "short" or "long." Adrenaline uses this classification to boost long requests from the start. Unfortunately, Adrenaline does not easily generalize to other LC services, whose features will be different. Additionally, a downside of request classification is that it cannot rank requests within each category, and therefore, is not fine-grained enough to accurately pinpoint requests at the tail; instead, an entire class of requests are boosted when only the longest ones needed to. To address this issue, ReTail instead follows the latency-based approach described next.
• Latency-based methods take a step further over classification, by using latency prediction to guide power management. Recent work [29,51] has shown great improvement over classification-based methods. Rubik [29] estimates a latency distribution for each request based on the current queue length
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- DynamoLLM: Designing LLM Inference Clusters for Performance and Energy EfficiencyJovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas 等HPCA 2025 · 被引用 106 次
- Harmonizing Efficiency and Practicability: Optimizing Resource Utilization in Serverless Computing with JiaguQingyuan Liu, Yanning Yang, Dong Du, Yubin Xia 等USENIX ATC 2024 · 被引用 39 次
- EcoFaaS: Rethinking the Design of Serverless Environments for Energy EfficiencyJovan Stojkovic, Nikoleta Iliakopoulou, Tianyin Xu, Hubertus Franke 等ISCA 2024 · 被引用 29 次
- throttLL'eM: Predictive GPU Throttling for Energy Efficient LLM Inference ServingAndreas Kosmas Kakolyris, Dimosthenis Masouros, Petros Vavaroutsos, Sotirios Xydis 等HPCA 2025 · 被引用 15 次
- SmartOClock: Workload- and Risk-Aware Overclocking in the CloudJovan Stojkovic, Pulkit A. Misra, Íñigo Goiri, Sam Whitlock 等ISCA 2024 · 被引用 12 次
它引用的顶会 Paper2
相关 Paper
- ANT-man: towards agile power management in the microservice eraXiaofeng Hou, Chao Li, Jiacheng Liu, Lu Zhang 等SC 2020 · 被引用 35 次
- Twig: Multi-Agent Task Management for Colocated Latency-Critical Cloud ServicesRajiv Nishtala, Vinicius Petrucci, Paul M. Carpenter, Magnus SjälanderHPCA 2020 · 被引用 76 次
- DDPC: Automated Data-Driven Power-Performance Controller Design on-the-fly for Latency-sensitive Web ServicesMehmet Savasci, Ahmed Ali-Eldin, Johan Eker, Anders Robertsson 等WWW 2023 · 被引用 6 次
- NMAP: Power Management Based on Network Packet Processing Mode Transition for Latency-Critical WorkloadsKi-Dong Kang, Gyeongseo Park, Hyosang Kim, Mohammad Alian 等MICRO 2021 · 被引用 18 次
- Intelligent Resource Scheduling for Co-located Latency-critical Services: A Multi-Model Collaborative Learning ApproachLei Liu, Xinglei Dou, Yuetao ChenFAST 2023
