ReGate: Enabling Power Gating in Neural Processing Units
Yuqi Xue, Jian Huang
Abstract
The energy efficiency of neural processing units (NPU) plays a critical role in developing sustainable data centers. Our study with different generations of NPU chips reveals that 30%-72% of their energy consumption is contributed by static power dissipation, due to the lack of power management support in modern NPU chips.
In this paper, we present ReGate, which enables fine-grained power-gating of each hardware component in NPU chips with hardware/software co-design. Unlike conventional power-gating techniques for generic processors, enabling power-gating in NPUs faces unique challenges due to the fundamental difference in hardware architecture and program execution model. To address these challenges, we carefully investigate the power-gating opportunities in each component of NPU chips and decide the best-fit power management scheme (i.e., hardware-vs. software-managed power gating). Specifically, for systolic arrays (SAs) that have deterministic execution patterns, ReGate enables cycle-level power gating at the granularity of processing elements (PEs) following the inherent dataflow execution in SAs. For inter-chip interconnect (ICI) and HBM controllers that have long idle intervals, ReGate employs a lightweight hardware-based idle-detection mechanism. For vector units and SRAM whose idle periods vary significantly depending on workload patterns, ReGate extends the NPU ISA and allows software (e.g., compilers) to manage the power gating. With implementation on a production-level NPU simulator, we show that ReGate can reduce the energy consumption of NPU chips by up to 32.8% (15.5% on average), with negligible impact on AI workload performance. The hardware implementation of power-gating logic introduces less than 3.3% overhead in NPU chips.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c0e50fca-8063-4e68-8b56-37977e03782dCited by top-tier papers1
Ask how each one uses itBuilds on19
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein et al.ASPLOS 2024 · 693 citations
- MLPerf Inference BenchmarkVijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson et al.ISCA 2020 · 517 citations
- ACT: designing sustainable computer systems with an architectural carbon modeling toolUdit Gupta, Mariam Elgamal, Gage Hills, Gu-Yeon Wei et al.ISCA 2022 · 176 citations
- AccelWattch: A Power Modeling Framework for Modern GPUsVijay Kandiah, Scott Peverelle, Mahmoud Khairy, Junrui Pan et al.MICRO 2021 · 134 citations
Related papers
- Exploiting Zero Data to Reduce Register File and Execution Unit Dynamic Power Consumption in GPGPUsAhmad M. Radaideh, Paul V. GratzDAC 2020 · 4 citations
- UPTPU: Improving Energy Efficiency of a Tensor Processing Unit through Underutilization Based Power-GatingPramesh Pandey, Noel Daniel Gundi, Koushik Chakraborty, Sanghamitra RoyDAC 2021 · 10 citations
- Hardware-Assisted Virtualization of Neural Processing Units for Cloud PlatformsYuqi Xue, Yiqi Liu, Lifeng Nai, Jian HuangMICRO 2024 · 10 citations
- PowerQuant: Architecture-Agnostic GPU Power Estimation via Quantile RegressionAditya Challa, Tanish Desai, Gargi Alavani Prabhu, Snehanshu Saha et al.HPDC 2026
- AgileWatts: An Energy-Efficient CPU Core Idle-State Architecture for Latency-Sensitive Server ApplicationsJawad Haj-Yahya, Haris Volos, Davide B. Bartolini, Georgia Antoniou et al.MICRO 2022 · 22 citations
