E.T.: re-thinking self-attention for transformer models on GPUs
Shiyang Chen, Shaoyi Huang, Santosh Pandey, Bingbing Li, Guang R. Gao, Long Zheng, Caiwen Ding, Hang Liu
摘要
Transformer-based deep learning models have become a ubiquitous vehicle to drive a variety of Natural Language Processing (NLP) related tasks beyond their accuracy ceiling. However, these models also suffer from two pronounced challenges, that is, gigantic model size and prolonged turnaround time. To this end, we introduce E.T. that rE-thinks self-attention computation for Transformer models on GPUs with the following contributions: First, we introduce a novel self-attention architecture, which encompasses two tailored self-attention operators with corresponding sequence length-aware optimizations, and operation reordering optimizations. Second, we present an attention-aware pruning design which judiciously uses various pruning algorithms to reduce more computations hence achieves significantly shorter turnaround time. For the pruning algorithms, we not only revamp the existing pruning algorithms, but also tailor new ones for transformer models. Taken together, we evaluate E.T. across a variety of benchmarks for Transformer, BERT BASE and DistilBERT, where E.T. presents superior performance over the mainstream projects, including the popular Nvidia Enterprise solutions, i.e., TensorRT and FasterTransformer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- A length adaptive algorithm-hardware co-design of transformer on FPGA through sparse attention and dynamic pipeliningHongwu Peng, Shaoyi Huang, Shiyang Chen, Bingbing Li 等DAC 2022 · 被引用 49 次
- TANGO: re-thinking quantization for graph neural network training on GPUsShiyang Chen, Da Zheng, Caiwen Ding, Chengying Huan 等SC 2023 · 被引用 10 次
- Dynamic Sparse Training via Balancing the Exploration-Exploitation Trade-offShaoyi Huang, Bowen Lei, Dongkuan Xu, Hongwu Peng 等DAC 2023 · 被引用 7 次
- Neurogenesis Dynamics-inspired Spiking Neural Network Training AccelerationShaoyi Huang, Haowen Fang, Kaleel Mahmood, Bowen Lei 等DAC 2023 · 被引用 6 次
- MaverIQ: Fingerprint-Guided Extrapolation and Fragmentation-Aware Layering for Intent-Based LLM ServingDimitrios Liakopoulos, Prasoon Sinha, Tianrui Hu, Myungjin Lee 等SC 2025 · 被引用 2 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu 等ICML 2020 · 被引用 1,773 次
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 被引用 695 次
相关 Paper
- DOTA: detect and omit weak attentions for scalable transformer accelerationZheng Qu, Liu Liu, Fengbin Tu, Zhaodong Chen 等ASPLOS 2022 · 被引用 131 次
- Dynamic N: M Fine-Grained Structured Sparse Attention MechanismZhaodong Chen, Zheng Qu, Yuying Quan, Liu Liu 等PPoPP 2023 · 被引用 26 次
- ELSA: Hardware-Software Co-design for Efficient, Lightweight Self-Attention Mechanism in Neural NetworksTae Jun Ham, Yejin Lee, Seong Hoon Seo, Soosung Kim 等ISCA 2021 · 被引用 185 次
- Accelerating attention through gradient-based learned runtime pruningZheng Li, Soroush Ghodrati, Amir Yazdanbakhsh, Hadi Esmaeilzadeh 等ISCA 2022 · 被引用 48 次
- DynaX: Sparse Attention Acceleration with Dynamic X: M Fine-Grained Structured PruningXiao Xiong, Zhaorui Chen, Yue Liang, Minghao Tian 等ASPLOS 2025 · 被引用 3 次
