Accurate and Convenient Energy Measurements for GPUs: A Detailed Study of NVIDIA GPU's Built-In Power Sensor
Zeyu Yang, Karel Adámek, Wesley Armour
摘要
The Graphical processing unit (GPU), initially developed for graphics rendering, has emerged as the go-to accelerator for high throughput and parallel workloads, spanning scientific simulations to AI, thanks to its performance and power efficiency. Given that 6 out of the top 10 fastest supercomputers in the world use NVIDIA GPUs and many AI companies each employ 10,000's of NVIDIA GPUs, an accurate understanding of GPU power consumption is essential for making progress to further improve its efficiency. Despite the limited documentation and the lack of understanding of its mechanisms, NVIDIA GPUs' built-in power sensor, providing easily accessible power readings via the nvidia-smi interface, is widely used in energy efficient computing research on GPUs. Our study seeks to elucidate the internal mechanisms of the power readings provided by nvidia-smi and assess the accuracy of the power and energy consumption data gathered from this method. We have developed a suite of micro-benchmarks to profile the behaviour of nvidia-smi power readings and have evaluated them on over 70 different GPUs from all architectural generations since power measurement was first introduced in the 'Fermi' generation of GPU. We have identified several unique and unforeseen problems in terms of power/energy measurement using nvidia-smi, for example on the A100 and H100 GPUs only 25% of the runtime is sampled for power consumption, during the other 75% of the time, the GPU can be using drastically different power and nvidia-smi and results presented by it are unaware of this. This along with other findings can lead to a drastic under/overestimation of energy consumed, especially when considering data centres housing tens of thousands of GPUs. We proposed several good practices that help to mitigate these problems. By comparing our results to those measured from an external power-meter, we have reduced the error in the energy measurement by an average of 35% and in some cases by as much as 65% in the test cases we present. We have encapsulated our learning, our micro-benchmark and measurement good practice into a Python library for public use.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Untangling GPU Power Consumption: Job-Level Inference in Cloud Shared SettingsPierre Jacquet, Maxime Agusti, Eddy Caron, Camille Coti 等EuroSys 2026 · 被引用 2 次
- The Hidden Joules: Evaluating the Energy Consumption of Vision Backbones for Progress Towards More Efficient Model InferenceZeyu Yang, Wesley ArmourICML 2025
- A Data-Centric Hardware Accelerator for Efficient Adaptive Radix TreeJin Zhao, Yu Zhang, Jun Huang, Weihang Yin 等DAC 2025
它引用的顶会 Paper2
相关 Paper
- Benchmark-driven Models for Energy Analysis and Attribution of GPU-Accelerated SupercomputingOscar Antepara, Zhengji Zhao, Brian Austin, Nan Ding 等SC 2025 · 被引用 5 次
- PowerQuant: Architecture-Agnostic GPU Power Estimation via Quantile RegressionAditya Challa, Tanish Desai, Gargi Alavani Prabhu, Snehanshu Saha 等HPDC 2026
- Characterizing Performance, Power, and Energy of AMD CDNA3 GPU FamilyBagus Hanindhito, Bhavesh PatelSC 2025 · 被引用 2 次
- Debunking the CUDA Myth Towards GPU-based AI Systems: Evaluation of the Performance and Programmability of Intel's Gaudi NPU for AI Model ServingYunjae Lee, Juntaek Lim, Jehyeon Bang, Eunyeong Cho 等ISCA 2025 · 被引用 2 次
- Dissecting and Modeling the Architecture of Modern GPU CoresRodrigo Huerta, Mojtaba Abaie Shoushtary, José-Lorenzo Cruz, Antonio GonzálezMICRO 2025 · 被引用 8 次
