DynaX: Sparse Attention Acceleration with Dynamic X: M Fine-Grained Structured Pruning
Xiao Xiong, Zhaorui Chen, Yue Liang, Minghao Tian, Jiaxing Shang, Jiang Zhong, Dajiang Liu
摘要
Owning to the mechanism of self-attention, Transformers have exhibited incredible performance in a wide range of artificial intelligence tasks. With the growth of sequence length, attention computation with quadratic complexity becomes the bottleneck, and dynamic sparsity is an effective technique to alleviate this problem. However, dynamic attention sparsity for long-sequence tasks suffers from two challenges, i.e., irregular sparse patterns and heavy prediction overhead. To this end, this paper proposes DynaX, an algorithm-hardware co-design framework that accelerates attention computation via dynamic X:M fine-grained structured pruning. Different from traditional N:M pruning, DynaX dynamically selects variable X (rather than a fixed N) important scores from a group via a 2-step pruning method, which results in high sparsity and less prediction memory overhead while maintaining pattern regularity to a certain extent. After that, DynaX performs block scheduling to reorganize score blocks into hardware blocks that can perfectly match the size of the processing element array (PEA), resulting in a higher utilization rate. Experimental results show that DynaX can achieve average sparsity of 89.54% and 91.77% for short-sequence tasks and long-sequence tasks, respectively, with less than 1% accuracy loss. Compared to Sanger and SALO2, DynaX achieves a speedup of 1.99X and 1.50X on the BERT-base model, and an energy efficiency improvement of 5.16X and 4.20X, respectively.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- LongSight: Compute-Enabled Memory to Accelerate Large-Context LLMs via Sparse AttentionDerrick Quinn, E. Ezgi Yücel, Jinkwon Kim, José F. Martínez 等MICRO 2025 · 被引用 3 次
- Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge ComputingTianhua Xia, Sai Qian ZhangMICRO 2025 · 被引用 2 次
相关 Paper
- A length adaptive algorithm-hardware co-design of transformer on FPGA through sparse attention and dynamic pipeliningHongwu Peng, Shaoyi Huang, Shiyang Chen, Bingbing Li 等DAC 2022 · 被引用 49 次
- Dynamic N: M Fine-Grained Structured Sparse Attention MechanismZhaodong Chen, Zheng Qu, Yuying Quan, Liu Liu 等PPoPP 2023 · 被引用 26 次
- FNM-Trans: Efficient FPGA-based Transformer Architecture with Full N: M SparsityManting Zhang, Jialin Cao, Kejia Shi, Keqing Zhao 等DAC 2024 · 被引用 10 次
- Efficient Transformer Inference with Statically Structured Sparse AttentionSteve Dai, Hasan Genc, Rangharajan Venkatesan, Brucek KhailanyDAC 2023 · 被引用 11 次
- SALO: an efficient spatial accelerator enabling hybrid sparse attention mechanisms for long sequencesGuan Shen, Jieru Zhao, Quan Chen, Jingwen Leng 等DAC 2022 · 被引用 37 次
