You Only Sample (Almost) Once: Linear Cost Self-Attention Via Bernoulli Sampling
Zhanpeng Zeng, Yunyang Xiong, Sathya N. Ravi, Shailesh Acharya, Glenn Moo Fung, Vikas Singh
摘要
Transformer-based models are widely used in natural language processing (NLP). Central to the transformer model is the self-attention mechanism, which captures the interactions of token pairs in the input sequences and depends quadratically on the sequence length. Training such models on longer sequences is expensive. In this paper, we show that a Bernoulli sampling attention mechanism based on Locality Sensitive Hashing (LSH), decreases the quadratic complexity of such models to linear. We bypass the quadratic cost by considering self-attention as a sum of individual tokens associated with Bernoulli random variables that can, in principle, be sampled at once by a single hash (although in practice, this number may be a small constant). This leads to an efficient sampling scheme to estimate self-attention which relies on specific modifications of LSH (to enable deployment on GPU architectures). We evaluate our algorithm on the GLUE benchmark with standard 512 sequence length where we see favorable performance relative to a standard pretrained Transformer. On the Long Range Arena (LRA) benchmark, for evaluating performance on long sequences, our method achieves results consistent with softmax self-attention but with sizable speed-ups and memory savings and often outperforms other efficient self-attention methods. Our code is available at https://github.com/mlpen/YOSO .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Flowformer: Linearizing Transformers with Conservation FlowsHaixu Wu, Jialong Wu, Jiehui Xu, Jianmin Wang 等ICML 2022 · 被引用 130 次
- Deep Unlearning via Randomized Conditionally Independent HessiansRonak Mehta, Sourav Pal, Vikas Singh, Sathya N. RaviCVPR 2022 · 被引用 50 次
- Primal-Attention: Self-attention through Asymmetric Kernel SVD in Primal RepresentationYingyi Chen, Qinghua Tao, Francesco Tonin, Johan A. K. SuykensNeurIPS 2023 · 被引用 42 次
- Graph Convolutions Enrich the Self-Attention in Transformers!Jeongwhan Choi, Hyowon Wi, Jayoung Kim, Yehjin Shin 等NeurIPS 2024 · 被引用 24 次
- Multi Resolution Analysis (MRA) for Approximate Self-AttentionZhanpeng Zeng, Sourav Pal, Jeffery Kline, Glenn Moo Fung 等ICML 2022 · 被引用 14 次
它引用的顶会 Paper7
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen 等ICLR 2021 · 被引用 881 次
- Nyströmformer: A Nyström-based Algorithm for Approximating Self-AttentionYunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan 等AAAI 2021 · 被引用 675 次
相关 Paper
- Sparse Attention with Learning to HashZhiqing Sun, Yiming Yang, Shinjae YooICLR 2022 · 被引用 21 次
- RACE Attention: A Strictly Linear-Time Attention for Long-Sequence TrainingSahil Joshi, Agniva Chowdhury, Amar Kanakamedala, Ekam Singh 等ICLR 2026 · 被引用 2 次
- Linear-Time Self Attention with Codeword Histogram for Efficient RecommendationYongji Wu, Defu Lian, Neil Zhenqiang Gong, Lu Yin 等WWW 2021 · 被引用 18 次
- SMYRF - Efficient Attention using Asymmetric ClusteringGiannis Daras, Nikita Kitaev, Augustus Odena, Alexandros G. DimakisNeurIPS 2020 · 被引用 54 次
- HyperAttention: Long-context Attention in Near-Linear TimeInsu Han, Rajesh Jayaram, Amin Karbasi, Vahab Mirrokni 等ICLR 2024 · 被引用 104 次
