AgilePkgC: An Agile System Idle State Architecture for Energy Proportional Datacenter Servers
Georgia Antoniou, Haris Volos, Davide B. Bartolini, Tom Rollet, Yiannakis Sazeides, Jawad Haj-Yahya
摘要
Modern user-facing applications deployed in datacenters use a distributed system architecture that exacerbates the latency requirements of their constituent microservices (30-250s). Existing CPU power-saving techniques degrade the performance of these applications due to the long transition latency (order of 100s) to wake up from a deep CPU idle state (C-state). For this reason, server vendors recommend only enabling shallow core C-states (e.g., CC1) for idle CPU cores, thus preventing the system from entering deep package C-states (e.g., PC6) when all CPU cores are idle. This choice, however, impairs server energy proportionality since power-hungry resources (e.g., IOs, uncore, DRAM) remain active even when there is no active core to use them. As we show, it is common for all cores to be idle due to the low average utilization (e.g., 5-20%) of datacenter servers running user-facing applications. We propose to reap this opportunity with AgilePkgC (APC), a new package C-state architecture that improves the energy proportionality of server processors running latency-critical applications. APC implements PC 1A (package C l agile), a new deep package C-state that a system can enter once all cores are in a shallow C-state (i.e., CC1) and has a nanosecond-scale transition latency. PC 1A is based on four key techniques. First, a hardware-based agile power management unit (APMU) rapidly detects when all cores enter a shallow core C-state (CC1) and triggers the system-level power savings control flow. Second, an IO Standby Mode (IOSM) places IO interfaces (e.g., PCIe, DMI, UPI, DRAM) in shallow (nanosecond-scale transition latency) low-power modes. Third, a CLM Retention (CLMR) mode rapidly reduces the CLM (Cache-and-home-agent, Last-level-cache, and Mesh network-on-chip) domain’s voltage to its retention level, drastically reducing its power consumption. Fourth, APC keeps all system PLLs active in PC 1A to allow nanosecond-scale exit latency by avoiding PLL re-locking overhead. Combining these techniques enables significant power savings while requiring less than 200ns transition latency, faster than existing deep package C-states (e.g., PC6), making PC 1A practical for datacenter servers. Our evaluation based on an Intel Skylake-based server shows that APC reduces the energy consumption of Memcached by up to 41% (25% on average) with <0.1% performance degradation. APC provides similar benefits for other representative workloads.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- AgileWatts: An Energy-Efficient CPU Core Idle-State Architecture for Latency-Sensitive Server ApplicationsJawad Haj-Yahya, Haris Volos, Davide B. Bartolini, Georgia Antoniou 等MICRO 2022 · 被引用 22 次
- BlitzCoin: Fully Decentralized Hardware Power Management for Accelerator-Rich SoCsMartin Cochet, Karthik Swaminathan, Erik Jens Loscalzo, Joseph Zuckerman 等ISCA 2024 · 被引用 7 次
它引用的顶会 Paper4
- A large scale analysis of hundreds of in-memory cache clusters at TwitterJuncheng Yang, Yao Yue, K. V. RashmiOSDI 2020 · 被引用 245 次
- IChannels: Exploiting Current Management Mechanisms to Create Covert Channels in Modern ProcessorsJawad Haj-Yahya, Lois Orosa, Jeremie S. Kim, Juan Gómez-Luna 等ISCA 2021 · 被引用 19 次
- NMAP: Power Management Based on Network Packet Processing Mode Transition for Latency-Critical WorkloadsKi-Dong Kang, Gyeongseo Park, Hyosang Kim, Mohammad Alian 等MICRO 2021 · 被引用 18 次
- FlexWatts: A Power- and Workload-Aware Hybrid Power Delivery Network for Energy-Efficient MicroprocessorsJawad Haj-Yahya, Mohammed Alser, Jeremie S. Kim, Lois Orosa 等MICRO 2020 · 被引用 12 次
相关 Paper
- GreenDIMM: OS-assisted DRAM Power Management for DRAM with a Sub-array Granularity Power-Down StateSeunghak Lee, Ki-Dong Kang, Hwanjun Lee, Hyungwon Park 等MICRO 2021 · 被引用 14 次
- EcoCore: Dynamic Core Management for Improving Energy Efficiency in Latency-Critical ApplicationsGyeongseo Park, Minho Kim, Ki-Dong Kang, Yunhyeong Jeon 等MICRO 2025 · 被引用 1 次
- Agile-DRAM: Agile Trade-Offs in Memory Capacity, Latency, and Energy for Data CentersJaeyoon Lee, Wonyeong Jung, Dongwhee Kim, Daero Kim 等HPCA 2024 · 被引用 4 次
- When Idling is Ideal: Optimizing Tail-Latency for Heavy-Tailed Datacenter Workloads with PerséphoneHenri Maxime Demoulin, Joshua Fried, Isaac Pedisich, Marios Kogias 等SOSP 2021 · 被引用 39 次
- TiNA: Tiered Network Buffer Architecture for Fast Networking in Chiplet-based CPUsSiddharth Agarwal, Tianchen Wang, Jinghan Huang, Saksham Agarwal 等ASPLOS 2026 · 被引用 1 次
