Lune

HPCA2022Top-tier venue

ReTail: Opting for Learning Simplicity to Enable QoS-Aware Power Management in the Cloud

Shuang Chen, Angela Jin, Christina Delimitrou, José F. Martínez

2022Year
34Citations
9Top-tier citations

Abstract

Many cloud services have Quality-of-Service (QoS) requirements; most requests have to to complete within a given latency constraint. Recently, researchers have begun to investigate whether it is possible to meet QoS while attempting to save power on a per-request basis. Existing work shows that one can indeed hand-tune a request latency predictor offline for a particular cloud application, and consult it at runtime to modulate CPU voltage and frequency, resulting in substantial power savings.

In this paper, we propose ReTail, an automated and general solution for request-level power management of latency-critical services with QoS constraints. We present a systematic process to select the features of any given application that best correlate with its request latency. ReTail uses these features to predict latency, and adjust CPU's power consumption. ReTail's predictor is trained fully at runtime. We show that unlike previous findings, simple techniques perform better than complex machine learning models, when using the right input features. For a web search engine, ReTail outperforms prior mechanisms based on complex hand-tuned predictors for that application domain. Furthermore, ReTail's systematic approach also yields superior power savings across a diverse set of cloud applications. TABLE I: Qualitative comparison of ReTail vs. two closely related proposals for power management of latency-critical applications. Rubik Gemini ReTail Method Statistical model Neural net Linear regression Feature space N/A Request Request&Application Feature selection N/A Hand-picked Systematic Request-accurate General applications Training overhead Low High Low Inference overhead Low High Low QoS guarantee No dropped requests Adapt to model drift dynamically adjust to system or application interference by retraining the latency predictor online. II. RELATED WORK Most proposals towards improving power efficiency [15, 26, 43] focus on throughput-oriented batch jobs. Recent work that explores DVFS for LC services with latency/QoS constraints can be broadly classified into two categories: Coarse-grained power management like Pegasus [34] dynamically adjusts frequency for the entire LC application. It addresses long-term latency variations under load fluctuations, but leaves power savings on the table by not differentiating individual requests.

Fine-grained power management outperforms coarsegrained management by differentiating at request granularity, generally boosting long and slowing down short requests, while trying to meet QoS. There are two main methods:

• Classification-based methods classify requests into short and long using various metrics, and boost all requests in the long category. EETL [23] tracks the progress of each request; those that exceed a predetermined execution threshold are flagged as long requests, and EETL boosts the frequency at that point. However, by the time a request reaches the progress threshold, it may be too late to prevent tail latency degradation. Adrenaline [24] introduces the concept of featuredriven request classification. It studies two LC applications, web search and key-value stores, and identifies, by human inspection, request features in each that can be used to predict incoming requests as being either "short" or "long." Adrenaline uses this classification to boost long requests from the start. Unfortunately, Adrenaline does not easily generalize to other LC services, whose features will be different. Additionally, a downside of request classification is that it cannot rank requests within each category, and therefore, is not fine-grained enough to accurately pinpoint requests at the tail; instead, an entire class of requests are boosted when only the longest ones needed to. To address this issue, ReTail instead follows the latency-based approach described next.

• Latency-based methods take a step further over classification, by using latency prediction to guide power management. Recent work [29,51] has shown great improvement over classification-based methods. Rubik [29] estimates a latency distribution for each request based on the current queue length

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext bfe68959-caf8-4758-b7c7-3c6614d48bdc

Cited by top-tier papers9

Ask how each one uses it

Builds on2

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines