SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
Hanrui Wang, Zhekai Zhang, Song Han
摘要
The attention mechanism is becoming increasingly popular in Natural Language Processing (NLP) applications, showing superior performance than convolutional and recurrent architectures. However, attention becomes the computation bottleneck because of its quadratic computational complexity to input length, complicated data movement and low arithmetic intensity. Moreover, existing NN accelerators mainly focus on optimizing convolutional or recurrent models, and cannot efficiently support attention. In this paper, we present SpAtten, an efficient algorithm-architecture co-design that leverages token sparsity, head sparsity, and quantization opportunities to reduce the attention computation and memory access. Inspired by the high redundancy of human languages, we propose the novel cascade token pruning to prune away unimportant tokens in the sentence. We also propose cascade head pruning to remove unessential heads. Cascade pruning is fundamentally different from weight pruning since there is no trainable weight in the attention mechanism, and the pruned tokens and heads are selected on the fly. To efficiently support them on hardware, we design a novel top-k engine to rank token and head importance scores with high throughput. Furthermore, we propose progressive quantization that first fetches MSBs only and performs the computation; if the confidence is low, it fetches LSBs and recomputes the attention outputs, trading computation for memory reduction.
Extensive experiments on 30 benchmarks show that, on average, SpAtten reduces DRAM access by 10.0× with no accuracy loss, and achieves 1.6×, 3.0×, 162×, 347× speedup, and 1.4×, 3.2×, 1193×, 4059× energy savings over A 3 accelerator, MNNFast accelerator, TITAN Xp GPU, Xeon CPU, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper111
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen 等NeurIPS 2023 · 被引用 1,003 次
- Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-FreeZihan Qiu, Zekun Wang, Bo Zheng, Zeyu Huang 等NeurIPS 2025 · 被引用 336 次
- A Fast Post-Training Pruning Framework for TransformersWoosuk Kwon, Sehoon Kim, Michael W. Mahoney, Joseph Hassoun 等NeurIPS 2022 · 被引用 247 次
- QuantumNAS: Noise-Adaptive Search for Robust Quantum CircuitsHanrui Wang, Yongshan Ding, Jiaqi Gu, Yujun Lin 等HPCA 2022 · 被引用 199 次
- S3: Increasing GPU Utilization during Generative Inference for Higher ThroughputYunho Jin, Chun-Feng Wu, David Brooks, Gu-Yeon WeiNeurIPS 2023 · 被引用 150 次
它引用的顶会 Paper40
- SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN TrainingEric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella 等HPCA 2020 · 被引用 490 次
- HyGCN: A GCN Accelerator with Hybrid ArchitectureMingyu Yan, Lei Deng, Xing Hu, Ling Liang 等HPCA 2020 · 被引用 338 次
- GCN-RL Circuit Designer: Transferable Transistor Sizing with Graph Neural Networks and Reinforcement LearningHanrui Wang, Kuan Wang, Jiacheng Yang, Linxiao Shen 等DAC 2020 · 被引用 326 次
- SpArch: Efficient Architecture for Sparse Matrix MultiplicationZhekai Zhang, Hanrui Wang, Song Han, William J. DallyHPCA 2020 · 被引用 280 次
- A3: Accelerating Attention Mechanisms in Neural Networks with ApproximationTae Jun Ham, Sungjun Jung, Seonghak Kim, Young H. Oh 等HPCA 2020 · 被引用 241 次
相关 Paper
- Blaze: An Efficient Bit-Sparse Attention Architecture With Workload Orchestration OptimizationRunzhou Zhang, Faxian Sun, Yiming Wang, Kunchen Zou 等DAC 2025
- Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token SelectionDongwon Jo, Beomseok Kang, Jiwon Song, jae-joon kimICML 2026 · 被引用 1 次
- DynaX: Sparse Attention Acceleration with Dynamic X: M Fine-Grained Structured PruningXiao Xiong, Zhaorui Chen, Yue Liang, Minghao Tian 等ASPLOS 2025 · 被引用 3 次
- CTA: Hardware-Software Co-design for Compressed Token Attention MechanismHaoran Wang, Haobo Xu, Ying Wang, Yinhe HanHPCA 2023 · 被引用 18 次
- Libra: A Hybrid-Sparse Attention Accelerator Featuring Multi-Level Workload BalanceFaxian Sun, Runzhou Zhang, Zhenyu Liu, Heng Liao 等DAC 2025 · 被引用 1 次
