FNM-Trans: Efficient FPGA-based Transformer Architecture with Full N: M Sparsity
Manting Zhang, Jialin Cao, Kejia Shi, Keqing Zhao, Genhao Zhang, Jun Yu, Kun Wang
摘要
Transformer models have become popular in various AI applications due to their exceptional performance. However, their impressive performance comes with significant computing and memory costs, hindering efficient deployment of Transformer-based applications. Many solutions focus on leveraging sparsity in weight matrix and attention computation. However, previous studies fail to exploit unified sparse pattern to accelerate all three modules of Transformer (QKV generation, attention computation and FFN). In this paper, we propose FNM-Trans, an adaptable and efficient algorithm-hardware co-design aimed at optimizing all three modules of the Transformer by fully harnessing N : M sparsity. At the algorithm level, we fully explore the interplay of dynamic pruning with static pruning under high N : M sparsity. At the hardware level, we develop a dedicated hardware architecture featuring a custom computing engine and a softmax module, tailored to support varying levels of N : M sparsity. Experiment results show that, our algorithm optimizes accuracy by 11.03% under 2:16 attention sparsity and 4:16 weight sparsity, compared to other methods. Additionally, FNM-Trans achieves speedups of 27.13× and 21.24× over Intel i9-9900X and NVIDIA RTX 2080 Ti, respectively, and outpaces current FPGA-based Transformers by 1.88× to 36.51×.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- FACT: FFN-Attention Co-optimized Transformer Architecture with Eager Correlation PredictionYubin Qin, Yang Wang, Dazheng Deng, Zhiren Zhao 等ISCA 2023 · 被引用 113 次
- Adaptable Butterfly Accelerator for Attention-based NNs via Hardware and Algorithm Co-designHongxiang Fan, Thomas Chau, Stylianos I. Venieris, Royson Lee 等MICRO 2022 · 被引用 63 次
- A length adaptive algorithm-hardware co-design of transformer on FPGA through sparse attention and dynamic pipeliningHongwu Peng, Shaoyi Huang, Shiyang Chen, Bingbing Li 等DAC 2022 · 被引用 49 次
- DynaX: Sparse Attention Acceleration with Dynamic X: M Fine-Grained Structured PruningXiao Xiong, Zhaorui Chen, Yue Liang, Minghao Tian 等ASPLOS 2025 · 被引用 3 次
- FLAME: Fully Leveraging MoE Sparsity for Transformer on FPGAXuanda Lin, Huinan Tian, Wenxiao Xue, Lanqi Ma 等DAC 2024 · 被引用 9 次
