Fast Monte-Carlo Approximation of the Attention Mechanism
Hyunjun Kim, JeongGil Ko
Abstract
We introduce Monte-Carlo Attention (MCA), a randomized approximation method for reducing the computational cost of self-attention mechanisms in Transformer architectures. MCA exploits the fact that the importance of each token in an input sequence vary with respect to their attention scores; thus, some degree of error can be tolerable when encoding tokens with low attention. Using approximate matrix multiplication, MCA applies different error bounds to encode input tokens such that those with low attention scores are computed with relaxed precision, whereas errors of salient elements are minimized. MCA can operate in parallel with other attention optimization schemes and does not require model modification. We study the theoretical error bounds and demonstrate that MCA reduces attention complexity (in FLOPS) for various Transformer models by up to 11 in GLUE benchmarks without compromising model accuracy. Source code and appendix: https://github.com/eis-lab/monte-carlo-attention
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 08e1e635-d240-41b4-9ae7-db600f87ee6fCited by top-tier papers1
Ask how each one uses itBuilds on9
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERTSheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma et al.AAAI 2020 · 656 citations
- Synthesizer: Rethinking Self-Attention for Transformer ModelsYi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan et al.ICML 2021 · 399 citations
- HAT: Hardware-Aware Transformers for Efficient Natural Language ProcessingHanrui Wang, Zhanghao Wu, Zhijian Liu, Han Cai et al.ACL 2020 · 215 citations
- BERT-of-Theseus: Compressing BERT by Progressive Module ReplacingCanwen Xu, Wangchunshu Zhou, Tao Ge, Furu Wei et al.EMNLP 2020 · 168 citations
Related papers
- Fine- and Coarse-Granularity Hybrid Self-Attention for Efficient BERTJing Zhao, Yifan Wang, Junwei Bao, Youzheng Wu et al.ACL 2022 · 7 citations
- MonarchAttention: Zero-Shot Conversion to Fast, Hardware-Aware Structured AttentionCan Yaras, Alec S. Xu, Pierre Abillama, Changwoo Lee et al.NeurIPS 2025 · 5 citations
- Multi Resolution Analysis (MRA) for Approximate Self-AttentionZhanpeng Zeng, Sourav Pal, Jeffery Kline, Glenn Moo Fung et al.ICML 2022 · 14 citations
- Token Statistics Transformer: Linear-Time Attention via Variational Rate ReductionZiyang Wu, Tianjiao Ding, Yifu Lu, Druv Pai et al.ICLR 2025
- Fast Transformers with Clustered AttentionApoorv Vyas, Angelos Katharopoulos, François FleuretNeurIPS 2020 · 193 citations
