Efficient Transformer Inference with Statically Structured Sparse Attention
Steve Dai, Hasan Genc, Rangharajan Venkatesan, Brucek Khailany
摘要
Self-attention matrices of Transformers are often highly sparse because the relevant context of each token is typically limited to just a few other tokens in the sequence. To reduce the computational burden of self-attention on Transformer inference, we propose static, structured, sparse attention masks that split attention matrices into dense regions, skipping computations outside these regions while reducing computations inside these regions. To support the proposed mask structure, we design an entropy-aware finetuning algorithm to naturally encourage attention sparsity while maximizing task accuracy. Furthermore, we extend a typical dense deep learning accelerator to efficiently exploit our structured sparsity pattern. Compared to a dense baseline, we achieve 56.6% reduction in energy consumption, 58.9% performance improvement with <1% accuracy loss and 2.6% area overhead.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper7
- ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV CachingYoupeng Zhao, Di Wu, Jun WangISCA 2024 · 被引用 35 次
- ESPACE: Dimensionality Reduction of Activations for Model CompressionCharbel Sakr, Brucek KhailanyNeurIPS 2024 · 被引用 21 次
- CHAI: Clustered Head Attention for Efficient LLM InferenceSaurabh Agarwal, Bilge Acun, Basil Hosmer, Mostafa Elhoushi 等ICML 2024 · 被引用 16 次
- Re-ttention: Ultra Sparse Visual Generation via Attention Statistical ReshapeRuichen Chen, Keith G. Mills, Liyao Jiang, Chao Gao 等NeurIPS 2025 · 被引用 10 次
- Growing a Twig to Accelerate Large Vision-Language ModelsZhenwei Shao, Mingyang Wang, Zhou Yu, Wenwen Pan 等ICCV 2025 · 被引用 3 次
相关 Paper
- SpARC: Token Similarity-Aware Sparse Attention Transformer Accelerator via Row-wise ClusteringHan Cho, Dongjun Kim, Seung-Eon Hwang, Jongsun ParkDAC 2024 · 被引用 7 次
- SWAT: Scalable and Efficient Window Attention-based Transformers Acceleration on FPGAsZhenyu Bai, Pranav Dangi, Huize Li, Tulika MitraDAC 2024 · 被引用 12 次
- DynaX: Sparse Attention Acceleration with Dynamic X: M Fine-Grained Structured PruningXiao Xiong, Zhaorui Chen, Yue Liang, Minghao Tian 等ASPLOS 2025 · 被引用 3 次
- SSFT: Algorithm and Hardware Co-design for Structured Sparse Fine-Tuning of Large Language ModelsMiao Yu, Trevor E. CarlsonDAC 2025 · 被引用 1 次
- SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated TilingHuizheng Wang, Jiahao Fang, Xinru Tang, Zhiheng Yue 等MICRO 2024 · 被引用 31 次
