GPNPU: Enabling Efficient Hardware-Based Direct Convolution with Multi-Precision Support in GPU Tensor Cores
Zhuoran Song, Jianfei Wang, Tianjian Li, Li Jiang, Jing Ke, Xiaoyao Liang, Naifeng Jing
Abstract
To tailor for DNN (Deep Neural Network) acceleration, GPU has migrated to new architectures such as NVIDIA Volta and Turing that incorporate dedicated Tensor Cores. Although good at GEMM (generic matrix-matrix multiplication), Tensor Cores still have inefficiency facing convolutions with certain layer structures. This paper proposes a GPNPU (General-Purpose Neural-network Processing Unit) architecture, which offers another option of direct convolution in GPU. It stitches the direct convolution dataflow into the Tensor Cores with little hardware support, and resorts to regulated data layout with stripe-mined convolution execution to achieve higher performance and power efficiency, while retaining the general programability as GPU. We further apply a unified core design to support varied operand types and precision for higher computing throughput. The evaluation shows that GPNPU can outperform Tensor Cores on typical DNNs by 1.4X for inference (FP16) and 1.2X for training with much reduced power. The INT8 performance even increases to 2.4X. Our study demonstrates that it is possible and appealing to refine the Tensor Cores for greater DNN acceleration, while conforming to GPU architecture for the programmability necessary in future DNN evolution.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 0e2b9c09-e3a3-4867-93af-8cb972631f47Cited by top-tier papers3
- DASP: Specific Dense Matrix Multiply-Accumulate Units Accelerated General Sparse Matrix-Vector MultiplicationYuechen Lu, Weifeng LiuSC 2023 · 37 citations
- GTuner: tuning DNN computations on GPU via graph attention networkQi Sun, Xinyun Zhang, Hao Geng, Yuxuan Zhao et al.DAC 2022 · 10 citations
- Characterizing Matrix Multiplication Units across General Parallel Patterns in Scientific ComputingYuechen Lu, Hongwei Zeng, Marc Casas, Weifeng LiuPPoPP 2026 · 1 citation
Related papers
- Duplo: Lifting Redundant Memory Accesses of Deep Neural Networks for GPU Tensor CoresHyeonjin Kim, Sungwoo Ahn, Yunho Oh, Bogil Kim et al.MICRO 2020 · 27 citations
- Cooperative Warp Execution in Tensor Core for RISC-V GPGPUAbubakr Nada, Giuseppe Maria Sarda, Erwan LenormandHPCA 2025 · 3 citations
- Optimizing batched Winograd convolution on GPUsDa Yan, Wei Wang, Xiaowen ChuPPoPP 2020 · 64 citations
- RM-STC: Row-Merge Dataflow Inspired GPU Sparse Tensor Core for Energy-Efficient Sparse AccelerationGuyue Huang, Zhengyang Wang, Po-An Tsai, Chen Zhang et al.MICRO 2023 · 15 citations
- Balancing Efficiency and Flexibility for DNN Acceleration via Temporal GPU-Systolic Array IntegrationCong Guo, Yangjie Zhou, Jingwen Leng, Yuhao Zhu et al.DAC 2020 · 35 citations
