Exploring the Performance Improvement of Tensor Processing Engines through Transformation in the Bit-weight Dimension of MACs
Qizhe Wu, Huawen Liang, Yuchen Gui, Zhichen Zeng, Zerong He, Linfeng Tao, Xiaotian Wang, Letian Zhao, Zhaoxi Zeng, Wei Yuan, Wei Wu, Xi Jin
摘要
General matrix-matrix multiplication (GEMM), serving as a cornerstone of AI computations, has positioned tensor processing engines (TPEs) as increasingly critical components within existing GPUs and domain-specific architectures (DSA). Our analysis identifies that the prevailing architectures primarily focus on dataflow or operand reuse strategies, when considering the combination of matrix multiplication with multiply-accumulator (MAC) itself, it provides greater optimization space for the design of TPEs. This work introduces a novel perspective on matrix multiplication from a hardware standpoint, focusing on the bit-weight dimension of MACs. Through this lens, we propose a finer-grained TPE notation, using matrix triple loops as an example, introducing new methods and ideas for designing and optimizing PE microarchitecture. Based on the new notation and transformations, we propose four optimization techniques that achieve varying degrees of improvement in timing, area, and power consumption. We implement our design in RTL using the SMIC-28nm process. Applying our methods to four classic TPE architectures (include systolic array [20], 3D-Cube [27], multiplier-adder tree [48], and 2D-Matrix [30]), we achieved area efficiency improvements of , and , and , and for energy efficiency respectively. When applied to a bit-slice architecture, we achieved a improvement in energy efficiency and in area efficiency compared to Laconic [38]. Our Verilog HDL code, along with timing, area, and power reports for circuit synthesis in URL: https://github.com/wqzustc/High-Performance-Tensor-Processing-Engines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Interstellar: Using Halide's Scheduling Language to Analyze DNN AcceleratorsXuan Yang, Mingyu Gao, Qiaoyi Liu, Jeff Setter 等ASPLOS 2020 · 被引用 237 次
- Distilling Bit-level Sparsity Parallelism for General Purpose Deep Learning AccelerationHang Lu, Liang Chang, Chenglong Li, Zixuan Zhu 等MICRO 2021 · 被引用 54 次
- BitWave: Exploiting Column-Based Bit-Level Sparsity for Deep Learning AccelerationMan Shi, Vikram Jain, Antony Joseph, Maurice Meijer 等HPCA 2024 · 被引用 46 次
- Trapezoid: A Versatile Accelerator for Dense and Sparse Matrix MultiplicationsYifan Yang, Joel S. Emer, Daniel SánchezISCA 2024 · 被引用 38 次
- TileFlow: A Framework for Modeling Fusion Dataflow via Tree-based AnalysisSize Zheng, Siyuan Chen, Siyuan Gao, Liancheng Jia 等MICRO 2023 · 被引用 31 次
相关 Paper
- MHE-TPE: Multi-Operand High-Radix Encoder for Mixed-Precision Fixed-Point Tensor Processing EnginesQizhe Wu, Jinyi Zhou, Zhanhe Hu, Zhichen Zeng 等MICRO 2025 · 被引用 2 次
- Duplo: Lifting Redundant Memory Accesses of Deep Neural Networks for GPU Tensor CoresHyeonjin Kim, Sungwoo Ahn, Yunho Oh, Bogil Kim 等MICRO 2020 · 被引用 27 次
- Virgo: Cluster-level Matrix Unit Integration in GPUs for Scalability and Energy EfficiencyHansung Kim, Ruohan Richard Yan, Joshua You, Tieliang Vamber Yang 等ASPLOS 2025 · 被引用 1 次
- Empowering Vector Architectures for ML: The CAMP Architecture for Matrix MultiplicationMohammadreza Esmali Nojehdeh, Hossein Mokhtarnia, Julian Pavon, Narcís Rodas 等MICRO 2025 · 被引用 1 次
- VSpGEMM: Exploiting Versal ACAP for High-Performance SpGEMM AccelerationKai Shi, Zhe Lin, Xinya Luan, Jianwang Zhai 等DAC 2025 · 被引用 1 次
