Gemini: Learning to Manage CPU Power for Latency-Critical Search Engines
Liang Zhou, Laxmi N. Bhuyan, K. K. Ramakrishnan
Abstract
Saving energy for latency-critical applications like web search can be challenging because of their strict tail latency constraints. State-of-the-art power management frameworks use Dynamic Voltage and Frequency Scaling (DVFS) and Sleep states techniques to slow down the request processing and finish the search just-in-time. However, accurately predicting the compute demand of a request can be difficult. In this paper, we present Gemini, a novel power management framework for latency-critical search engines. Gemini has two unique features to capture the per query service time variation. First, at light loads without request queuing, a two-step DVFS is used to manage the CPU power. Our two-step DVFS selects the initial CPU frequency based on the query specific service time prediction and then judiciously boosts the initial frequency at the right time to catch-up to the deadline. The determination of boosting time further relies on estimating the error in the prediction of individual query's service time. At high loads, where there is request queuing, only the current request being executed and the critical request in the queue adopt a two-step DVFS. All the other requests in-between use the same frequency to reduce the frequency transition overhead. Second, we develop two separate neural network models, one for predicting the service time and the other for the error in the prediction. The combination of these two predictors significantly improves the power saving and tail latency results of our two-step DVFS. Gemini is implemented on the Solr search engine. Evaluations on three representative query traces show that Gemini saves 41% of the CPU power, and is better than other state-of-the-art techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af1812f1-58b5-4119-9761-6284334f3e3aCited by top-tier papers8
- DynamoLLM: Designing LLM Inference Clusters for Performance and Energy EfficiencyJovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas et al.HPCA 2025 · 106 citations
- ReTail: Opting for Learning Simplicity to Enable QoS-Aware Power Management in the CloudShuang Chen, Angela Jin, Christina Delimitrou, José F. MartínezHPCA 2022 · 34 citations
- EcoFaaS: Rethinking the Design of Serverless Environments for Energy EfficiencyJovan Stojkovic, Nikoleta Iliakopoulou, Tianyin Xu, Hubertus Franke et al.ISCA 2024 · 29 citations
- CoolEdge: hotspot-relievable warm water cooling for energy-efficient edge datacentersQiangyu Pei, Shutong Chen, Qixia Zhang, Xinhui Zhu et al.ASPLOS 2022 · 28 citations
- SmartOClock: Workload- and Risk-Aware Overclocking in the CloudJovan Stojkovic, Pulkit A. Misra, Íñigo Goiri, Sam Whitlock et al.ISCA 2024 · 12 citations
Related papers
- Cottage: Coordinated Time Budget Assignment for Latency, Quality and Power Optimization in Web SearchLiang Zhou, Laxmi N. Bhuyan, K. K. RamakrishnanHPCA 2022 · 2 citations
- NMAP: Power Management Based on Network Packet Processing Mode Transition for Latency-Critical WorkloadsKi-Dong Kang, Gyeongseo Park, Hyosang Kim, Mohammad Alian et al.MICRO 2021 · 18 citations
- AgileWatts: An Energy-Efficient CPU Core Idle-State Architecture for Latency-Sensitive Server ApplicationsJawad Haj-Yahya, Haris Volos, Davide B. Bartolini, Georgia Antoniou et al.MICRO 2022 · 22 citations
- Prediction-Informed Power Management for General-Purpose Compute ServersJonggyu Park, Simon Peter, Thomas E. AndersonEuroSys 2026
- PowerGrad: Hierarchical Power Management for Power-Limited ML Inference ClustersHyoungwook Nam, Raghavendra Pradyumna Pothukuchi, Alper Buyuktosunoglu, Aporva Amarnath et al.ISCA 2026 · 1 citation
