Predict; Don't React for Enabling Efficient Fine-Grain DVFS in GPUs
Srikant Bharadwaj, Shomit Das, Kaushik Mazumdar, Bradford M. Beckmann, Stephen Kosonocky
摘要
With the continuous improvement of on-chip integrated voltage regulators (IVRs) and fast, adaptive frequency control, dynamic voltage-frequency scaling (DVFS) transition times have shrunk from the microsecond to the nanosecond regime, providing additional opportunities to improve energy efficiency. The key to unlocking the continued improvement in V/f circuit technology is the creation of new, smarter DVFS mechanisms that better adapt to rapid fluctuations in workload demand.
It is particularly important to optimize fine-grain DVFS mechanisms for graphics processing units (GPUs) as the chips become ever more important workhorses in the datacenter. However, GPU's massive amount of thread-level parallelism makes it uniquely difficult to determine the optimal V/f state at run-time. Existing solutions-mostly designed for singlethreaded CPUs and longer time scales-fail to consider the seemingly chaotic, highly varying nature of GPU workloads at short time scales.
This paper proposes a novel prediction mechanism, PCSTALL, that is tailored for emerging DVFS capabilities in GPUs and achieves near-optimal energy efficiency. Using the insights from our fine-grained workload analysis, we propose a wavefront-level program counter (PC) based DVFS mechanism that improves program behavior prediction accuracy by 32% on average for a wide set of GPU applications at 1µs DVFS time epochs. Compared to the current state-of-art, our PC-based technique achieves 19% average improvement when optimized for Energy-Delay 2 Product (ED 2 P) at 50µs time epochs, reaching 32% power efficiencies when operated with 1µs DVFS technologies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Using Analytical Performance/Power Model and Fine-Grained DVFS to Enhance AI Accelerator Energy EfficiencyZibo Wang, Yijia Zhang, Fuchun Wei, Bingqiang Wang 等ASPLOS 2025 · 被引用 9 次
- HALO: Hardware-Aware Quantization with Low Critical-Path-Delay Weights for LLM AccelerationRohan Juneja, Shivam Aggarwal, Safeen Huda, Tulika Mitra 等AAAI 2026 · 被引用 2 次
- Beyond Utilization: Energy-Conscious GPU Sharing for Inference ServingPrasoon Sinha, Dimitrios Liakopoulos, Nathan Lemma, Neeraja J. YadwadkarSOSP 2026
它引用的顶会 Paper2
相关 Paper
- Improving GPU Energy Efficiency through an Application-transparent Frequency Scaling Policy with Performance AssuranceYijia Zhang, Qiang Wang, Zhe Lin, Pengxiang Xu 等EuroSys 2024 · 被引用 18 次
- Know Your Enemy To Save Cloud Energy: Energy-Performance Characterization of Machine Learning ServingJunyeol Yu, Jongseok Kim, Euiseong SeoHPCA 2023 · 被引用 14 次
- A Workload-Aware DVFS Robust to Concurrent Tasks for Mobile DevicesChengdong Lin, Kun Wang, Zhenjiang Li, Yu PuMobiCom 2023 · 被引用 52 次
- Warped-Compaction: Maximizing GPU Register File Bandwidth Utilization via Operand CompactionEunbi Jeong, Ipoom Jeong, Myung Kuk Yoon, Nam Sung KimHPCA 2025 · 被引用 2 次
- PowerWeave: Unlocking Energy-Efficient ML on GPUs with OS-Level Spatial Power ManagementVasilis Kypriotis, Eric Dubberstein, Patrick H. Coppock, Eliot H. Solomon 等ISCA 2026 · 被引用 1 次
