You Only Sample (Almost) Once: Linear Cost Self-Attention Via Bernoulli Sampling
Zhanpeng Zeng, Yunyang Xiong, Sathya N. Ravi, Shailesh Acharya, Glenn Moo Fung, Vikas Singh
Abstract
Transformer-based models are widely used in natural language processing (NLP). Central to the transformer model is the self-attention mechanism, which captures the interactions of token pairs in the input sequences and depends quadratically on the sequence length. Training such models on longer sequences is expensive. In this paper, we show that a Bernoulli sampling attention mechanism based on Locality Sensitive Hashing (LSH), decreases the quadratic complexity of such models to linear. We bypass the quadratic cost by considering self-attention as a sum of individual tokens associated with Bernoulli random variables that can, in principle, be sampled at once by a single hash (although in practice, this number may be a small constant). This leads to an efficient sampling scheme to estimate self-attention which relies on specific modifications of LSH (to enable deployment on GPU architectures). We evaluate our algorithm on the GLUE benchmark with standard 512 sequence length where we see favorable performance relative to a standard pretrained Transformer. On the Long Range Arena (LRA) benchmark, for evaluating performance on long sequences, our method achieves results consistent with softmax self-attention but with sizable speed-ups and memory savings and often outperforms other efficient self-attention methods. Our code is available at https://github.com/mlpen/YOSO .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 56cabeff-c2c3-4a49-a524-1f45e5a456e2Cited by top-tier papers14
- Flowformer: Linearizing Transformers with Conservation FlowsHaixu Wu, Jialong Wu, Jiehui Xu, Jianmin Wang et al.ICML 2022 · 130 citations
- Deep Unlearning via Randomized Conditionally Independent HessiansRonak Mehta, Sourav Pal, Vikas Singh, Sathya N. RaviCVPR 2022 · 50 citations
- Primal-Attention: Self-attention through Asymmetric Kernel SVD in Primal RepresentationYingyi Chen, Qinghua Tao, Francesco Tonin, Johan A. K. SuykensNeurIPS 2023 · 42 citations
- Graph Convolutions Enrich the Self-Attention in Transformers!Jeongwhan Choi, Hyowon Wi, Jayoung Kim, Yehjin Shin et al.NeurIPS 2024 · 24 citations
- Multi Resolution Analysis (MRA) for Approximate Self-AttentionZhanpeng Zeng, Sourav Pal, Jeffery Kline, Glenn Moo Fung et al.ICML 2022 · 14 citations
Builds on7
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen et al.ICLR 2021 · 881 citations
- Nyströmformer: A Nyström-based Algorithm for Approximating Self-AttentionYunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan et al.AAAI 2021 · 675 citations
Related papers
- Sparse Attention with Learning to HashZhiqing Sun, Yiming Yang, Shinjae YooICLR 2022 · 21 citations
- RACE Attention: A Strictly Linear-Time Attention for Long-Sequence TrainingSahil Joshi, Agniva Chowdhury, Amar Kanakamedala, Ekam Singh et al.ICLR 2026 · 2 citations
- Linear-Time Self Attention with Codeword Histogram for Efficient RecommendationYongji Wu, Defu Lian, Neil Zhenqiang Gong, Lu Yin et al.WWW 2021 · 18 citations
- SMYRF - Efficient Attention using Asymmetric ClusteringGiannis Daras, Nikita Kitaev, Augustus Odena, Alexandros G. DimakisNeurIPS 2020 · 54 citations
- HyperAttention: Long-context Attention in Near-Linear TimeInsu Han, Rajesh Jayaram, Amin Karbasi, Vahab Mirrokni et al.ICLR 2024 · 104 citations
