SC2021Top-tier venue
E.T.: re-thinking self-attention for transformer models on GPUs
Shiyang Chen, Shaoyi Huang, Santosh Pandey, Bingbing Li, Guang R. Gao, Long Zheng, Caiwen Ding, Hang Liu
Abstract
Transformer-based deep learning models have become a ubiquitous vehicle to drive a variety of Natural Language Processing (NLP) related tasks beyond their accuracy ceiling. However, these models also suffer from two pronounced challenges, that is, gigantic model size and prolonged turnaround time. To this end, we introduce E.T. that rE-thinks self-attention computation for Transformer models on GPUs with the following contributions: First, we introduce a novel self-attention architecture, which encompasses two tailored self-attention operators with corresponding sequence length-aware optimizations, and operation reordering optimizations. Second, we present an attention-aware pruning design which judiciously uses various pruning algorithms to reduce more computations hence achieves significantly shorter turnaround time. For the pruning algorithms, we not only revamp the existing pruning algorithms, but also tailor new ones for transformer models. Taken together, we evaluate E.T. across a variety of benchmarks for Transformer, BERT BASE and DistilBERT, where E.T. presents superior performance over the mainstream projects, including the popular Nvidia Enterprise solutions, i.e., TensorRT and FasterTransformer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 88e3af93-57e2-424c-a431-eee25d01a27eCited by top-tier papers6
- A length adaptive algorithm-hardware co-design of transformer on FPGA through sparse attention and dynamic pipeliningHongwu Peng, Shaoyi Huang, Shiyang Chen, Bingbing Li et al.DAC 2022 · 49 citations
- TANGO: re-thinking quantization for graph neural network training on GPUsShiyang Chen, Da Zheng, Caiwen Ding, Chengying Huan et al.SC 2023 · 10 citations
- Dynamic Sparse Training via Balancing the Exploration-Exploitation Trade-offShaoyi Huang, Bowen Lei, Dongkuan Xu, Hongwu Peng et al.DAC 2023 · 7 citations
- Neurogenesis Dynamics-inspired Spiking Neural Network Training AccelerationShaoyi Huang, Haowen Fang, Kaleel Mahmood, Bowen Lei et al.DAC 2023 · 6 citations
- MaverIQ: Fingerprint-Guided Extrapolation and Fragmentation-Aware Layering for Intent-Based LLM ServingDimitrios Liakopoulos, Prasoon Sinha, Tianrui Hu, Myungjin Lee et al.SC 2025 · 2 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 695 citations
Related papers
- DOTA: detect and omit weak attentions for scalable transformer accelerationZheng Qu, Liu Liu, Fengbin Tu, Zhaodong Chen et al.ASPLOS 2022 · 131 citations
- Dynamic N: M Fine-Grained Structured Sparse Attention MechanismZhaodong Chen, Zheng Qu, Yuying Quan, Liu Liu et al.PPoPP 2023 · 26 citations
- ELSA: Hardware-Software Co-design for Efficient, Lightweight Self-Attention Mechanism in Neural NetworksTae Jun Ham, Yejin Lee, Seong Hoon Seo, Soosung Kim et al.ISCA 2021 · 185 citations
- Accelerating attention through gradient-based learned runtime pruningZheng Li, Soroush Ghodrati, Amir Yazdanbakhsh, Hadi Esmaeilzadeh et al.ISCA 2022 · 48 citations
- DynaX: Sparse Attention Acceleration with Dynamic X: M Fine-Grained Structured PruningXiao Xiong, Zhaorui Chen, Yue Liang, Minghao Tian et al.ASPLOS 2025 · 3 citations
