ReTail: Opting for Learning Simplicity to Enable QoS-Aware Power Management in the Cloud
Shuang Chen, Angela Jin, Christina Delimitrou, José F. Martínez
Abstract
Many cloud services have Quality-of-Service (QoS) requirements; most requests have to to complete within a given latency constraint. Recently, researchers have begun to investigate whether it is possible to meet QoS while attempting to save power on a per-request basis. Existing work shows that one can indeed hand-tune a request latency predictor offline for a particular cloud application, and consult it at runtime to modulate CPU voltage and frequency, resulting in substantial power savings.
In this paper, we propose ReTail, an automated and general solution for request-level power management of latency-critical services with QoS constraints. We present a systematic process to select the features of any given application that best correlate with its request latency. ReTail uses these features to predict latency, and adjust CPU's power consumption. ReTail's predictor is trained fully at runtime. We show that unlike previous findings, simple techniques perform better than complex machine learning models, when using the right input features. For a web search engine, ReTail outperforms prior mechanisms based on complex hand-tuned predictors for that application domain. Furthermore, ReTail's systematic approach also yields superior power savings across a diverse set of cloud applications. TABLE I: Qualitative comparison of ReTail vs. two closely related proposals for power management of latency-critical applications. Rubik Gemini ReTail Method Statistical model Neural net Linear regression Feature space N/A Request Request&Application Feature selection N/A Hand-picked Systematic Request-accurate General applications Training overhead Low High Low Inference overhead Low High Low QoS guarantee No dropped requests Adapt to model drift dynamically adjust to system or application interference by retraining the latency predictor online. II. RELATED WORK Most proposals towards improving power efficiency [15, 26, 43] focus on throughput-oriented batch jobs. Recent work that explores DVFS for LC services with latency/QoS constraints can be broadly classified into two categories: Coarse-grained power management like Pegasus [34] dynamically adjusts frequency for the entire LC application. It addresses long-term latency variations under load fluctuations, but leaves power savings on the table by not differentiating individual requests.
Fine-grained power management outperforms coarsegrained management by differentiating at request granularity, generally boosting long and slowing down short requests, while trying to meet QoS. There are two main methods:
• Classification-based methods classify requests into short and long using various metrics, and boost all requests in the long category. EETL [23] tracks the progress of each request; those that exceed a predetermined execution threshold are flagged as long requests, and EETL boosts the frequency at that point. However, by the time a request reaches the progress threshold, it may be too late to prevent tail latency degradation. Adrenaline [24] introduces the concept of featuredriven request classification. It studies two LC applications, web search and key-value stores, and identifies, by human inspection, request features in each that can be used to predict incoming requests as being either "short" or "long." Adrenaline uses this classification to boost long requests from the start. Unfortunately, Adrenaline does not easily generalize to other LC services, whose features will be different. Additionally, a downside of request classification is that it cannot rank requests within each category, and therefore, is not fine-grained enough to accurately pinpoint requests at the tail; instead, an entire class of requests are boosted when only the longest ones needed to. To address this issue, ReTail instead follows the latency-based approach described next.
• Latency-based methods take a step further over classification, by using latency prediction to guide power management. Recent work [29,51] has shown great improvement over classification-based methods. Rubik [29] estimates a latency distribution for each request based on the current queue length
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bfe68959-caf8-4758-b7c7-3c6614d48bdcCited by top-tier papers9
- DynamoLLM: Designing LLM Inference Clusters for Performance and Energy EfficiencyJovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas et al.HPCA 2025 · 106 citations
- Harmonizing Efficiency and Practicability: Optimizing Resource Utilization in Serverless Computing with JiaguQingyuan Liu, Yanning Yang, Dong Du, Yubin Xia et al.USENIX ATC 2024 · 39 citations
- EcoFaaS: Rethinking the Design of Serverless Environments for Energy EfficiencyJovan Stojkovic, Nikoleta Iliakopoulou, Tianyin Xu, Hubertus Franke et al.ISCA 2024 · 29 citations
- throttLL'eM: Predictive GPU Throttling for Energy Efficient LLM Inference ServingAndreas Kosmas Kakolyris, Dimosthenis Masouros, Petros Vavaroutsos, Sotirios Xydis et al.HPCA 2025 · 15 citations
- SmartOClock: Workload- and Risk-Aware Overclocking in the CloudJovan Stojkovic, Pulkit A. Misra, Íñigo Goiri, Sam Whitlock et al.ISCA 2024 · 12 citations
Builds on2
- Accelerometer: Understanding Acceleration Opportunities for Data Center Overheads at HyperscaleAkshitha Sriraman, Abhishek DhanotiaASPLOS 2020 · 78 citations
- Gemini: Learning to Manage CPU Power for Latency-Critical Search EnginesLiang Zhou, Laxmi N. Bhuyan, K. K. RamakrishnanMICRO 2020 · 27 citations
Related papers
- ANT-man: towards agile power management in the microservice eraXiaofeng Hou, Chao Li, Jiacheng Liu, Lu Zhang et al.SC 2020 · 35 citations
- Twig: Multi-Agent Task Management for Colocated Latency-Critical Cloud ServicesRajiv Nishtala, Vinicius Petrucci, Paul M. Carpenter, Magnus SjälanderHPCA 2020 · 76 citations
- DDPC: Automated Data-Driven Power-Performance Controller Design on-the-fly for Latency-sensitive Web ServicesMehmet Savasci, Ahmed Ali-Eldin, Johan Eker, Anders Robertsson et al.WWW 2023 · 6 citations
- NMAP: Power Management Based on Network Packet Processing Mode Transition for Latency-Critical WorkloadsKi-Dong Kang, Gyeongseo Park, Hyosang Kim, Mohammad Alian et al.MICRO 2021 · 18 citations
- Intelligent Resource Scheduling for Co-located Latency-critical Services: A Multi-Model Collaborative Learning ApproachLei Liu, Xinglei Dou, Yuetao ChenFAST 2023
