Fine- and Coarse-Granularity Hybrid Self-Attention for Efficient BERT
Jing Zhao, Yifan Wang, Junwei Bao, Youzheng Wu, Xiaodong He
Abstract
Transformer-based pre-trained models, such as BERT, have shown extraordinary success in achieving state-of-the-art results in many natural language processing applications. However, deploying these models can be prohibitively costly, as the standard self-attention mechanism of the Transformer suffers from quadratic computational cost in the input sequence length. To confront this, we propose FCA, a fine- and coarse-granularity hybrid self-attention that reduces the computation cost through progressively shortening the computational sequence length in self-attention. Specifically, FCA conducts an attention-based scoring strategy to determine the informativeness of tokens at each layer. Then, the informative tokens serve as the fine-granularity computing units in self-attention and the uninformative tokens are replaced with one or several clusters as the coarse-granularity computing units in self-attention. Experiments on the standard GLUE benchmark show that BERT with FCA achieves 2x reduction in FLOPs over original BERT with <1% loss in accuracy. We show that FCA offers a significantly better trade-off between accuracy and FLOPs compared to prior methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e4385786-09fc-458e-a028-eba963935b7cCited by top-tier papers1
Ask how each one uses itBuilds on11
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 695 citations
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu et al.ACL 2020 · 660 citations
- A Unified MRC Framework for Named Entity RecognitionXiaoya Li, Jingrong Feng, Yuxian Meng, Qinghong Han et al.ACL 2020 · 617 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- The Lottery Ticket Hypothesis for Pre-trained BERT NetworksTianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu et al.NeurIPS 2020 · 428 citations
Related papers
- Fast Transformers with Clustered AttentionApoorv Vyas, Angelos Katharopoulos, François FleuretNeurIPS 2020 · 193 citations
- Pyramid-BERT: Reducing Complexity via Successive Core-set based Token SelectionXin Huang, Ashish Khetan, Rene Bidart, Zohar S. KarninACL 2022
- Attribution-Driven Adaptive Token Pruning for TransformersYaoyao Yan, Hui Yu, Weizhi XuNeurIPS 2025 · 1 citation
- Token Dropping for Efficient BERT PretrainingLe Hou, Richard Yuanzhe Pang, Tianyi Zhou, Yuexin Wu et al.ACL 2022
- Fast Monte-Carlo Approximation of the Attention MechanismHyunjun Kim, JeongGil KoAAAI 2022 · 8 citations
