A length adaptive algorithm-hardware co-design of transformer on FPGA through sparse attention and dynamic pipelining
Hongwu Peng, Shaoyi Huang, Shiyang Chen, Bingbing Li, Tong Geng, Ang Li, Weiwen Jiang, Wujie Wen, Jinbo Bi, Hang Liu, Caiwen Ding
摘要
Transformers are considered one of the most important deep learning models since 2018, in part because it establishes state-of-the-art (SOTA) records and could potentially replace existing Deep Neural Networks (DNNs). Despite the remarkable triumphs, the prolonged turnaround time of Transformer models is a widely recognized roadblock. The variety of sequence lengths imposes additional computing overhead where inputs need to be zero-padded to the maximum sentence length in the batch to accommodate the parallel computing platforms. This paper targets the field-programmable gate array (FPGA) and proposes a coherent sequence length adaptive algorithm-hardware co-design for Transformer acceleration. Particularly, we develop a hardware-friendly sparse attention operator and a length-aware hardware resource scheduling algorithm. The proposed sparse attention operator brings the complexity of attention-based models down to linear complexity and alleviates the off-chip memory traffic. The proposed length-aware resource hardware scheduling algorithm dynamically allocates the hardware resources to fill up the pipeline slots and eliminates bubbles for NLP tasks. Experiments show that our design has very small accuracy loss and has 80.2 × and 2.6 × speedup compared to CPU and GPU implementation, and 4 × higher energy efficiency than state-of-the-art GPU accelerator optimized via CUBLAS GEMM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- LinGCN: Structural Linearized Graph Convolutional Network for Homomorphically Encrypted InferenceHongwu Peng, Ran Ran, Yukui Luo, Jiahui Zhao 等NeurIPS 2023 · 被引用 57 次
- AutoReP: Automatic ReLU Replacement for Fast Private Network InferenceHongwu Peng, Shaoyi Huang, Tong Zhou, Yukui Luo 等ICCV 2023 · 被引用 44 次
- Ising-Traffic: Using Ising Machine Learning to Predict Traffic Congestion under UncertaintyZhenyu Pan, Anshujit Sharma, Jerry Yao-Chieh Hu, Zhuo Liu 等AAAI 2023 · 被引用 43 次
- Visual Prompting Upgrades Neural Network Sparsification: A Data-Model PerspectiveCan Jin, Tianjin Huang, Yihua Zhang, Mykola Pechenizkiy 等AAAI 2025 · 被引用 30 次
- Dynamic Sparse Training via Balancing the Exploration-Exploitation Trade-offShaoyi Huang, Bowen Lei, Dongkuan Xu, Hongwu Peng 等DAC 2023 · 被引用 7 次
它引用的顶会 Paper10
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- On the Relationship between Self-Attention and Convolutional LayersJean-Baptiste Cordonnier, Andreas Loukas, Martin JaggiICLR 2020 · 被引用 629 次
- SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head PruningHanrui Wang, Zhekai Zhang, Song HanHPCA 2021 · 被引用 412 次
相关 Paper
- Adaptable Butterfly Accelerator for Attention-based NNs via Hardware and Algorithm Co-designHongxiang Fan, Thomas Chau, Stylianos I. Venieris, Royson Lee 等MICRO 2022 · 被引用 63 次
- DOTA: detect and omit weak attentions for scalable transformer accelerationZheng Qu, Liu Liu, Fengbin Tu, Zhaodong Chen 等ASPLOS 2022 · 被引用 131 次
- DynaX: Sparse Attention Acceleration with Dynamic X: M Fine-Grained Structured PruningXiao Xiong, Zhaorui Chen, Yue Liang, Minghao Tian 等ASPLOS 2025 · 被引用 3 次
- FNM-Trans: Efficient FPGA-based Transformer Architecture with Full N: M SparsityManting Zhang, Jialin Cao, Kejia Shi, Keqing Zhao 等DAC 2024 · 被引用 10 次
- SWAT: Scalable and Efficient Window Attention-based Transformers Acceleration on FPGAsZhenyu Bai, Pranav Dangi, Huize Li, Tulika MitraDAC 2024 · 被引用 12 次
