Accelerating attention through gradient-based learned runtime pruning
Zheng Li, Soroush Ghodrati, Amir Yazdanbakhsh, Hadi Esmaeilzadeh, Mingu Kang
Abstract
Self-attention is a key enabler of state-of-art accuracy for various transformer-based Natural Language Processing models. This attention mechanism calculates a correlation score for each word with respect to the other words in a sentence. Commonly, only a small subset of words highly correlates with the word under attention, which is only determined at runtime. As such, a significant amount of computation is inconsequential due to low attention scores and can potentially be pruned. The main challenge is finding the threshold for the scores below which subsequent computation will be inconsequential. Although such a threshold is discrete, this paper formulates its search through a soft differentiable regularizer integrated into the loss function of the training. This formulation piggy backs on the back-propagation training to analytically co-optimize the threshold and the weights simultaneously, striking a formally optimal balance between accuracy and computation pruning. To best utilize this mathematical innovation, we devise a bit-serial architecture, dubbed LeOPArd, for transformer language models with bit-level early termination microarchitectural mechanism. We evaluate our design across 43 back-end tasks for MemN2N, BERT, ALBERT, GPT-2, and Vision transformer models. Post-layout results show that, on average, LeOPArd yields 1.9×and 3.9×speedup and energy reduction, respectively, while keeping the average accuracy virtually intact (< 0.2% degradation).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0ae40ac0-db03-4f4e-b527-08677f473970Cited by top-tier papers11
- FLAT: An Optimized Dataflow for Mitigating Attention BottlenecksSheng-Chun Kao, Suvinay Subramanian, Gaurav Agrawal, Amir Yazdanbakhsh et al.ASPLOS 2023 · 68 citations
- Sparse Attention Acceleration with Synergistic In-Memory Pruning and On-Chip RecomputationAmir Yazdanbakhsh, Ashkan Moradifirouzabadi, Zheng Li, Mingu KangMICRO 2022 · 47 citations
- Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous BatchingSungmin Yun, Kwanhee Kyung, Juhwan Cho, Jaewan Choi et al.MICRO 2024 · 40 citations
- SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated TilingHuizheng Wang, Jiahao Fang, Xinru Tang, Zhiheng Yue et al.MICRO 2024 · 31 citations
- Tandem Processor: Grappling with Emerging Operators in Neural NetworksSoroush Ghodrati, Sean Kinzer, Hanyang Xu, Rohan Mahapatra et al.ASPLOS 2024 · 21 citations
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
Related papers
- E.T.: re-thinking self-attention for transformer models on GPUsShiyang Chen, Shaoyi Huang, Santosh Pandey, Bingbing Li et al.SC 2021 · 13 citations
- EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP InferenceThierry Tambe, Coleman Hooper, Lillian Pentecost, Tianyu Jia et al.MICRO 2021 · 117 citations
- DOTA: detect and omit weak attentions for scalable transformer accelerationZheng Qu, Liu Liu, Fengbin Tu, Zhaodong Chen et al.ASPLOS 2022 · 131 citations
- DynaX: Sparse Attention Acceleration with Dynamic X: M Fine-Grained Structured PruningXiao Xiong, Zhaorui Chen, Yue Liang, Minghao Tian et al.ASPLOS 2025 · 3 citations
- SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head PruningHanrui Wang, Zhekai Zhang, Song HanHPCA 2021 · 412 citations
