FNM-Trans: Efficient FPGA-based Transformer Architecture with Full N: M Sparsity
Manting Zhang, Jialin Cao, Kejia Shi, Keqing Zhao, Genhao Zhang, Jun Yu, Kun Wang
Abstract
Transformer models have become popular in various AI applications due to their exceptional performance. However, their impressive performance comes with significant computing and memory costs, hindering efficient deployment of Transformer-based applications. Many solutions focus on leveraging sparsity in weight matrix and attention computation. However, previous studies fail to exploit unified sparse pattern to accelerate all three modules of Transformer (QKV generation, attention computation and FFN). In this paper, we propose FNM-Trans, an adaptable and efficient algorithm-hardware co-design aimed at optimizing all three modules of the Transformer by fully harnessing N : M sparsity. At the algorithm level, we fully explore the interplay of dynamic pruning with static pruning under high N : M sparsity. At the hardware level, we develop a dedicated hardware architecture featuring a custom computing engine and a softmax module, tailored to support varying levels of N : M sparsity. Experiment results show that, our algorithm optimizes accuracy by 11.03% under 2:16 attention sparsity and 4:16 weight sparsity, compared to other methods. Additionally, FNM-Trans achieves speedups of 27.13× and 21.24× over Intel i9-9900X and NVIDIA RTX 2080 Ti, respectively, and outpaces current FPGA-based Transformers by 1.88× to 36.51×.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- FACT: FFN-Attention Co-optimized Transformer Architecture with Eager Correlation PredictionYubin Qin, Yang Wang, Dazheng Deng, Zhiren Zhao et al.ISCA 2023 · 113 citations
- Adaptable Butterfly Accelerator for Attention-based NNs via Hardware and Algorithm Co-designHongxiang Fan, Thomas Chau, Stylianos I. Venieris, Royson Lee et al.MICRO 2022 · 63 citations
- A length adaptive algorithm-hardware co-design of transformer on FPGA through sparse attention and dynamic pipeliningHongwu Peng, Shaoyi Huang, Shiyang Chen, Bingbing Li et al.DAC 2022 · 49 citations
- DynaX: Sparse Attention Acceleration with Dynamic X: M Fine-Grained Structured PruningXiao Xiong, Zhaorui Chen, Yue Liang, Minghao Tian et al.ASPLOS 2025 · 3 citations
- FLAME: Fully Leveraging MoE Sparsity for Transformer on FPGAXuanda Lin, Huinan Tian, Wenxiao Xue, Lanqi Ma et al.DAC 2024 · 9 citations
