Predict; Don't React for Enabling Efficient Fine-Grain DVFS in GPUs
Srikant Bharadwaj, Shomit Das, Kaushik Mazumdar, Bradford M. Beckmann, Stephen Kosonocky
Abstract
With the continuous improvement of on-chip integrated voltage regulators (IVRs) and fast, adaptive frequency control, dynamic voltage-frequency scaling (DVFS) transition times have shrunk from the microsecond to the nanosecond regime, providing additional opportunities to improve energy efficiency. The key to unlocking the continued improvement in V/f circuit technology is the creation of new, smarter DVFS mechanisms that better adapt to rapid fluctuations in workload demand.
It is particularly important to optimize fine-grain DVFS mechanisms for graphics processing units (GPUs) as the chips become ever more important workhorses in the datacenter. However, GPU's massive amount of thread-level parallelism makes it uniquely difficult to determine the optimal V/f state at run-time. Existing solutions-mostly designed for singlethreaded CPUs and longer time scales-fail to consider the seemingly chaotic, highly varying nature of GPU workloads at short time scales.
This paper proposes a novel prediction mechanism, PCSTALL, that is tailored for emerging DVFS capabilities in GPUs and achieves near-optimal energy efficiency. Using the insights from our fine-grained workload analysis, we propose a wavefront-level program counter (PC) based DVFS mechanism that improves program behavior prediction accuracy by 32% on average for a wide set of GPU applications at 1µs DVFS time epochs. Compared to the current state-of-art, our PC-based technique achieves 19% average improvement when optimized for Energy-Delay 2 Product (ED 2 P) at 50µs time epochs, reaching 32% power efficiencies when operated with 1µs DVFS technologies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0f8629f5-920b-4f63-ace7-ea6b3a2c378bCited by top-tier papers3
- Using Analytical Performance/Power Model and Fine-Grained DVFS to Enhance AI Accelerator Energy EfficiencyZibo Wang, Yijia Zhang, Fuchun Wei, Bingqiang Wang et al.ASPLOS 2025 · 9 citations
- HALO: Hardware-Aware Quantization with Low Critical-Path-Delay Weights for LLM AccelerationRohan Juneja, Shivam Aggarwal, Safeen Huda, Tulika Mitra et al.AAAI 2026 · 2 citations
- Beyond Utilization: Energy-Conscious GPU Sharing for Inference ServingPrasoon Sinha, Dimitrios Liakopoulos, Nathan Lemma, Neeraja J. YadwadkarSOSP 2026
Builds on2
- Kite: A Family of Heterogeneous Interposer Topologies Enabled via Accurate Interconnect ModelingSrikant Bharadwaj, Jieming Yin, Bradford M. Beckmann, Tushar KrishnaDAC 2020 · 94 citations
- MDM: The GPU Memory Divergence ModelLu Wang, Magnus Jahre, Almutaz Adileh, Lieven EeckhoutMICRO 2020 · 27 citations
Related papers
- Improving GPU Energy Efficiency through an Application-transparent Frequency Scaling Policy with Performance AssuranceYijia Zhang, Qiang Wang, Zhe Lin, Pengxiang Xu et al.EuroSys 2024 · 18 citations
- Know Your Enemy To Save Cloud Energy: Energy-Performance Characterization of Machine Learning ServingJunyeol Yu, Jongseok Kim, Euiseong SeoHPCA 2023 · 14 citations
- A Workload-Aware DVFS Robust to Concurrent Tasks for Mobile DevicesChengdong Lin, Kun Wang, Zhenjiang Li, Yu PuMobiCom 2023 · 52 citations
- Warped-Compaction: Maximizing GPU Register File Bandwidth Utilization via Operand CompactionEunbi Jeong, Ipoom Jeong, Myung Kuk Yoon, Nam Sung KimHPCA 2025 · 2 citations
- PowerWeave: Unlocking Energy-Efficient ML on GPUs with OS-Level Spatial Power ManagementVasilis Kypriotis, Eric Dubberstein, Patrick H. Coppock, Eliot H. Solomon et al.ISCA 2026 · 1 citation
