SC2021Top-tier venue
Accelerating applications using edge tensor processing units
Kuan-Chieh Hsu, Hung-Wei Tseng
Abstract
Neural network (NN) accelerators have been integrated into a widespectrum of computer systems to accommodate the rapidly growing demands for artificial intelligence (AI) and machine learning (ML) applications. NN accelerators share the idea of providing native hardware support for operations on multidimensional tensor data. Therefore, NN accelerators are theoretically tensor processors that can improve system performance for any problem that uses tensors as inputs/outputs. Unfortunately, commercially available NN accelerators only expose computation capabilities through AI/MLspecific interfaces. Furthermore, NN accelerators reveal very few hardware design details, so applications cannot easily leverage the tensor operations NN accelerators provide.
This paper introduces General-Purpose Computing on Edge Tensor Processing Units (GPETPU), an open-source, open-architecture framework that allows the developer and research communities to discover opportunities that NN accelerators enable for applications. GPETPU includes a powerful programming interface with efficient runtime system-level support-similar to that of CUDA/OpenCL in GPGPU computing-to bridge the gap between application demands and mismatched hardware/software interfaces.
We built GPETPU machine uses Edge Tensor Processing Units (Edge TPUs), which are widely available and representative of many commercial NN accelerators. We identified several novel use cases and revisited the algorithms. By leveraging the underlying Edge TPUs to perform tensor-algorithm-based compute kernels, our results reveal that GPETPU can achieve a 2.46× speedup over highend CPUs and reduce energy consumption by 40%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- SIMD2: a generalized matrix instruction set for accelerating tensor computation beyond GEMMYunan Zhang, Po-An Tsai, Hung-Wei TsengISCA 2022 · 6 citations
- KAMI: Communication-Avoiding General Matrix Multiplication within a Single GPUHemeng Wang, Yang Du, Sidu Li, Xiaowen Tian et al.SC 2025 · 4 citations
- Characterizing Mobile SoC for Accelerating Heterogeneous LLM InferenceLe Chen, Dahu Feng, Erhu Feng, Yingrui Wang et al.SOSP 2025 · 3 citations
Builds on8
- SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN TrainingEric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella et al.HPCA 2020 · 490 citations
- HyGCN: A GCN Accelerator with Hybrid ArchitectureMingyu Yan, Lei Deng, Xing Hu, Ling Liang et al.HPCA 2020 · 338 citations
- SpArch: Efficient Architecture for Sparse Matrix MultiplicationZhekai Zhang, Hanrui Wang, Song Han, William J. DallyHPCA 2020 · 280 citations
- FlexTensor: An Automatic Schedule Exploration and Optimization Framework for Tensor Computation on Heterogeneous SystemSize Zheng, Yun Liang, Shuo Wang, Renze Chen et al.ASPLOS 2020 · 171 citations
- Tensaurus: A Versatile Accelerator for Mixed Sparse-Dense Tensor ComputationsNitish Kumar Srivastava, Hanchen Jin, Shaden Smith, Hongbo Rong et al.HPCA 2020 · 121 citations
Related papers
- TGraph: A Tensor-centric Graph Processing FrameworkYongliang Zhang, Yuanyuan Zhu, Hao Zhang, Congli Gao et al.SIGMOD 2025 · 1 citation
- PET: Optimizing Tensor Programs with Partially Equivalent Transformations and Automated CorrectionsHaojie Wang, Jidong Zhai, Mingyu Gao, Zixuan Ma et al.OSDI 2021 · 77 citations
- Tandem Processor: Grappling with Emerging Operators in Neural NetworksSoroush Ghodrati, Sean Kinzer, Hanyang Xu, Rohan Mahapatra et al.ASPLOS 2024 · 21 citations
- GPNPU: Enabling Efficient Hardware-Based Direct Convolution with Multi-Precision Support in GPU Tensor CoresZhuoran Song, Jianfei Wang, Tianjian Li, Li Jiang et al.DAC 2020 · 12 citations
- EdgeNN: Efficient Neural Network Inference for CPU-GPU Integrated Edge DevicesChenyang Zhang, Feng Zhang, Kuangyu Chen, Mingjun Chen et al.ICDE 2023 · 16 citations
