Efficient Attention via Control Variates
Lin Zheng, Jianbo Yuan, Chong Wang, Lingpeng Kong
Abstract
Random-feature-based attention (RFA) is an efficient approximation of softmax attention with linear runtime and space complexity. However, the approximation gap between RFA and conventional softmax attention is not well studied. Built upon previous progress of RFA, we characterize this gap through the lens of control variates and show that RFA can be decomposed into a sum of multiple control variate estimators for each element in the sequence. This new framework reveals that exact softmax attention can be recovered from RFA by manipulating each control variate. Besides, it allows us to develop a more flexible form of control variates, resulting in a novel attention mechanism that significantly reduces the approximation gap while maintaining linear complexity. Extensive experiments demonstrate that our model outperforms state-of-the-art efficient attention mechanisms on both vision and language tasks. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5c47de54-d01b-4879-81c8-0ed5ed23e684Cited by top-tier papers9
- Many-Shot In-Context LearningRishabh Agarwal, Avi Singh, Lei Zhang, Bernd Bohnet et al.NeurIPS 2024 · 271 citations
- Various Lengths, Constant Speed: Efficient Language Modeling with Lightning AttentionZhen Qin, Weigao Sun, Dong Li, Xuyang Shen et al.ICML 2024 · 26 citations
- MetaLA: Unified Optimal Linear Approximation to Softmax Attention MapYuhong Chou, Man Yao, Kexin Wang, Yuqi Pan et al.NeurIPS 2024 · 22 citations
- Coneheads: Hierarchy Aware AttentionAlbert Tseng, Tao Yu, Toni J. B. Liu, Christopher De SaNeurIPS 2023 · 8 citations
- Mobile Attention: Mobile-Friendly Linear-Attention for Vision TransformersZhiyu Yao, Jian Wang, Haixu Wu, Jingdong Wang et al.ICML 2024 · 6 citations
Builds on52
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
Related papers
- Random Feature AttentionHao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz et al.ICLR 2021 · 425 citations
- Linear Complexity Randomized Self-attention MechanismLin Zheng, Chong Wang, Lingpeng KongICML 2022 · 39 citations
- Dense-Exponential Random Features: Sharp Positive Estimators of the Gaussian KernelValerii Likhosherstov, Krzysztof Marcin Choromanski, Kumar Avinava Dubey, Frederick Liu et al.NeurIPS 2023 · 5 citations
- ELFATT: Efficient Linear Fast Attention for Vision TransformersChong Wu, Maolin Che, Renjie Xu, Zhuoheng Ran et al.ACM MM 2025 · 3 citations
- Breaking the Low-Rank Dilemma of Linear AttentionQihang Fan, Huaibo Huang, Ran HeCVPR 2025
