High Performance, Low Power Matrix Multiply Design on ACAP: from Architecture, Design Challenges and DSE Perspectives
Jinming Zhuang, Zhuoping Yang, Peipei Zhou
摘要
As the increasing complexity of Neural Network(NN) models leads to high demands for computation, AMD introduces a heterogeneous programmable system-on-chip (SoC), i.e., Versal ACAP architectures featured with programmable logic(PL), CPUs, and dedicated AI engines (AIE) ASICs which has a theoretical throughput up to 6.4 TFLOPs for FP32, 25.6 TOPs for INT16 and 102.4 TOPs for INT8. However, the higher level of complexity makes it non-trivial to achieve the theoretical performance even for well-studied applications like matrix-matrix multiply. In this paper, we provide AutoMM, an automatic white-box framework that can systematically generate the design for MM accelerators on Versal which achieves 3.7 TFLOPs, 7.5 TOPs, and 28.2 TOPs for FP32, INT16, and INT8 data type respectively. Our designs are tested on board and achieve gains of 7.20x (FP32), 3.26x (INT16), 6.23x (INT8) energy efficiency than AMD U250, 2.32x (FP32) than Nvidia Jetson TX2, 1.06x (FP32), 1.70x (INT8) than Nvidia A100.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper2
相关 Paper
- VSpGEMM: Exploiting Versal ACAP for High-Performance SpGEMM AccelerationKai Shi, Zhe Lin, Xinya Luan, Jianwang Zhai 等DAC 2025 · 被引用 1 次
- HeteroSVD: Efficient SVD Accelerator on Versal ACAP with Algorithm-Hardware Co-DesignXinya Luan, Zhe Lin, Kai Shi, Jianwang Zhai 等DAC 2025 · 被引用 1 次
- G2PM: Performance Modeling for ACAP Architecture with Dual-Tiered Graph Representation LearningTuo Dai, Bizhao Shi, Guojie LuoDAC 2024 · 被引用 1 次
- QiMeng-Xpiler: Transcompiling Tensor Programs for Deep Learning Systems with a Neural-Symbolic ApproachShouyang Dong, Jun Bi, Di Huang, Jiaming Guo 等OSDI 2025 · 被引用 6 次
- ADEPT: automatic differentiable DEsign of photonic tensor coresJiaqi Gu, Hanqing Zhu, Chenghao Feng, Zixuan Jiang 等DAC 2022 · 被引用 18 次
