Cost-Efficient Overclocking in Immersion-Cooled Datacenters
Majid Jalili, Ioannis Manousakis, Iñigo Goiri, Pulkit A. Misra, Ashish Raniwala, Husam Alissa, Bharath Ramakrishnan, Phillip Tuma, Christian Belady, Marcus Fontoura, Ricardo Bianchini
摘要
Cloud providers typically use air-based solutions for cooling servers in datacenters. However, increasing transistor counts and the end of Dennard scaling will result in chips with thermal design power that exceeds the capabilities of air cooling in the near future. Consequently, providers have started to explore liquid cooling solutions (e.g., cold plates, immersion cooling) for the most power-hungry workloads. By keeping the servers cooler, these new solutions enable providers to operate server components beyond the normal frequency range (i.e., overclocking them) all the time. Still, providers must tradeoff the increase in performance via overclocking with its higher power draw and any component reliability implications.
In this paper, we argue that two-phase immersion cooling (2PIC) is the most promising technology, and build three prototype 2PIC tanks. Given the benefits of 2PIC, we characterize the impact of overclocking on performance, power, and reliability. Moreover, we propose several new scenarios for taking advantage of overclocking in cloud platforms, including oversubscribing servers and virtual machine (VM) auto-scaling. For the autoscaling scenario, we build a system that leverages overclocking for either hiding the latency of VM creation or postponing the VM creations in the hopes of not needing them. Using realistic cloud workloads running on a tank prototype, we show that overclocking can improve performance by 20%, increase VM packing density by 20%, and improve tail latency in auto-scaling scenarios by 54%. The combination of 2PIC and overclocking can reduce platform cost by up to 13% compared to air cooling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Characterizing Power Management Opportunities for LLMs in the CloudPratyush Patel, Esha Choukse, Chaojie Zhang, Íñigo Goiri 等ASPLOS 2024 · 被引用 83 次
- Designing Cloud Servers for Lower CarbonJaylen Wang, Daniel S. Berger, Fiodar Kazhamiaka, Celine Irvene 等ISCA 2024 · 被引用 49 次
- TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud PlatformsJovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Esha Choukse 等ASPLOS 2025 · 被引用 30 次
- CoolEdge: hotspot-relievable warm water cooling for energy-efficient edge datacentersQiangyu Pei, Shutong Chen, Qixia Zhang, Xinhui Zhu 等ASPLOS 2022 · 被引用 28 次
- AgileWatts: An Energy-Efficient CPU Core Idle-State Architecture for Latency-Sensitive Server ApplicationsJawad Haj-Yahya, Haris Volos, Davide B. Bartolini, Georgia Antoniou 等MICRO 2022 · 被引用 22 次
它引用的顶会 Paper3
- Protean: VM Allocation Service at ScaleOri Hadary, Luke Marshall, Ishai Menache, Abhisek Pan 等OSDI 2020 · 被引用 189 次
- Prediction-Based Power Oversubscription in Cloud PlatformsAlok Gautam Kumbhare, Reza Azimi, Ioannis Manousakis, Anand Bonde 等USENIX ATC 2021 · 被引用 90 次
- Data Center Power Oversubscription with a Medium Voltage Power Plane and Priority-Aware CappingVarun Sakalkar, Vasileios Kontorinis, David Landhuis, Shaohong Li 等ASPLOS 2020 · 被引用 56 次
相关 Paper
- SmartOClock: Workload- and Risk-Aware Overclocking in the CloudJovan Stojkovic, Pulkit A. Misra, Íñigo Goiri, Sam Whitlock 等ISCA 2024 · 被引用 12 次
- CoINT2: A Heuristic Coordinator for Responsive Receive-Side Network I/O Virtualization in Overcommitment CloudXu Huan, Jian Li, Haibing GuanINFOCOM 2025 · 被引用 1 次
- Hyrax: Fail-in-Place Server Operation in Cloud PlatformsJialun Lyu, Marisa You, Celine Irvene, Mark Jung 等OSDI 2023 · 被引用 18 次
- You Can Always Get What You Want: CPU Virtualization Made Fast and FreeYun Wang, Xingguo Jia, Ben Luo, Kenan Liu 等SOSP 2026
- Thunderbolt: Throughput-Optimized, Quality-of-Service-Aware Power Capping at ScaleShaohong Li, Xi Wang, Xiao Zhang, Vasileios Kontorinis 等OSDI 2020 · 被引用 42 次
