UPTPU: Improving Energy Efficiency of a Tensor Processing Unit through Underutilization Based Power-Gating
Pramesh Pandey, Noel Daniel Gundi, Koushik Chakraborty, Sanghamitra Roy
Abstract
The AI boom is bringing a plethora of domain-specific architectures for Neural Network computations. Google’s Tensor Processing Unit (TPU), a Deep Neural Network (DNN) accelerator, has replaced the CPUs/GPUs in its data centers, claiming more than 15 × rate of inference. However, the unprecedented growth in DNN workloads with the widespread use of AI services projects an increasing energy consumption of TPU based data centers. In this work, we parametrize the extreme hardware underutilization in TPU systolic array and propose UPTPU: an intelligent, dataflow adaptive power-gating paradigm to provide a staggering 3.5 × – 6.5× energy efficiency to TPU for different input batch sizes.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 585d6469-9690-4e2d-bc5e-4b614486710cCited by top-tier papers2
- PowerPruning: Selecting Weights and Activations for Power-Efficient Neural Network AccelerationRichard Petri, Grace Li Zhang, Yiran Chen, Ulf Schlichtmann et al.DAC 2023 · 11 citations
- ReGate: Enabling Power Gating in Neural Processing UnitsYuqi Xue, Jian HuangMICRO 2025 · 8 citations
Related papers
- Partitioned Scheduling and Parallelism Assignment for Real-Time DNN Inference Tasks on Multi-TPUBinqi Sun, Tomasz Kloda, Chu-Ge Wu, Marco CaccamoDAC 2024 · 8 citations
- Know Your Enemy To Save Cloud Energy: Energy-Performance Characterization of Machine Learning ServingJunyeol Yu, Jongseok Kim, Euiseong SeoHPCA 2023 · 14 citations
- Dataflow Mirroring: Architectural Support for Highly Efficient Fine-Grained Spatial Multitasking on Systolic-Array NPUsJounghoo Lee, Jinwoo Choi, Jaeyeon Kim, Jinho Lee et al.DAC 2021 · 39 citations
- REDUCT: Keep it Close, Keep it Cool! : Efficient Scaling of DNN Inference on Multi-core CPUs with Near-Cache ComputeAnant V. Nori, Rahul Bera, Shankar Balachandran, Joydeep Rakshit et al.ISCA 2021 · 17 citations
- Accelerating applications using edge tensor processing unitsKuan-Chieh Hsu, Hung-Wei TsengSC 2021 · 33 citations
