Fast Monte-Carlo Approximation of the Attention Mechanism
Hyunjun Kim, JeongGil Ko
摘要
We introduce Monte-Carlo Attention (MCA), a randomized approximation method for reducing the computational cost of self-attention mechanisms in Transformer architectures. MCA exploits the fact that the importance of each token in an input sequence vary with respect to their attention scores; thus, some degree of error can be tolerable when encoding tokens with low attention. Using approximate matrix multiplication, MCA applies different error bounds to encode input tokens such that those with low attention scores are computed with relaxed precision, whereas errors of salient elements are minimized. MCA can operate in parallel with other attention optimization schemes and does not require model modification. We study the theoretical error bounds and demonstrate that MCA reduces attention complexity (in FLOPS) for various Transformer models by up to 11 in GLUE benchmarks without compromising model accuracy. Source code and appendix: https://github.com/eis-lab/monte-carlo-attention
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERTSheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma 等AAAI 2020 · 被引用 656 次
- Synthesizer: Rethinking Self-Attention for Transformer ModelsYi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan 等ICML 2021 · 被引用 399 次
- HAT: Hardware-Aware Transformers for Efficient Natural Language ProcessingHanrui Wang, Zhanghao Wu, Zhijian Liu, Han Cai 等ACL 2020 · 被引用 215 次
- BERT-of-Theseus: Compressing BERT by Progressive Module ReplacingCanwen Xu, Wangchunshu Zhou, Tao Ge, Furu Wei 等EMNLP 2020 · 被引用 168 次
相关 Paper
- Fine- and Coarse-Granularity Hybrid Self-Attention for Efficient BERTJing Zhao, Yifan Wang, Junwei Bao, Youzheng Wu 等ACL 2022 · 被引用 7 次
- MonarchAttention: Zero-Shot Conversion to Fast, Hardware-Aware Structured AttentionCan Yaras, Alec S. Xu, Pierre Abillama, Changwoo Lee 等NeurIPS 2025 · 被引用 5 次
- Multi Resolution Analysis (MRA) for Approximate Self-AttentionZhanpeng Zeng, Sourav Pal, Jeffery Kline, Glenn Moo Fung 等ICML 2022 · 被引用 14 次
- Token Statistics Transformer: Linear-Time Attention via Variational Rate ReductionZiyang Wu, Tianjiao Ding, Yifu Lu, Druv Pai 等ICLR 2025
- Fast Transformers with Clustered AttentionApoorv Vyas, Angelos Katharopoulos, François FleuretNeurIPS 2020 · 被引用 193 次
