Skyformer: Remodel Self-Attention with Gaussian Kernel and Nyström Method
Yifan Chen, Qi Zeng, Heng Ji, Yun Yang
摘要
Transformers are expensive to train due to the quadratic time and space complexity in the self-attention mechanism. On the other hand, although kernel machines suffer from the same computation bottleneck in pairwise dot products, several approximation schemes have been successfully incorporated to considerably reduce their computational cost without sacrificing too much accuracy. In this work, we leverage the computation methods for kernel machines to alleviate the high computational cost and introduce Skyformer, which replaces the softmax structure with a Gaussian kernel to stabilize the model training and adapts the Nyström method to a non-positive semidefinite matrix to accelerate the computation. We further conduct theoretical analysis by showing that the matrix approximation error of our proposed method is small in the spectral norm. Experiments on Long Range Arena benchmark show that the proposed method is sufficient in getting comparable or even better performance than the full self-attention while requiring fewer computation resources.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Simple linear attention language models balance the recall-throughput tradeoffSimran Arora, Sabri Eyuboglu, Michael Zhang, Aman Timalsina 等ICML 2024 · 被引用 154 次
- KERPLE: Kernelized Relative Positional Embedding for Length ExtrapolationTa-Chung Chi, Ting-Han Fan, Peter J. Ramadge, Alexander RudnickyNeurIPS 2022 · 被引用 112 次
- The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax MimicryMichael Zhang, Kush Bhatia, Hermann Kumbong, Christopher RéICLR 2024 · 被引用 103 次
- KDEformer: Accelerating Transformers via Kernel Density EstimationAmir Zandieh, Insu Han, Majid Daliri, Amin KarbasiICML 2023 · 被引用 55 次
- Uniform Memory Retrieval with Larger Capacity for Modern Hopfield ModelsDennis Wu, Jerry Yao-Chieh Hu, Teng-Yun Hsiao, Han LiuICML 2024 · 被引用 44 次
它引用的顶会 Paper12
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang 等AAAI 2021 · 被引用 7,289 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
相关 Paper
- cosFormer: Rethinking Softmax In AttentionZhen Qin, Weixuan Sun, Hui Deng, Dongxu Li 等ICLR 2022 · 被引用 303 次
- Nyströmformer: A Nyström-based Algorithm for Approximating Self-AttentionYunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan 等AAAI 2021 · 被引用 675 次
- Rethinking Attention with PerformersKrzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song 等ICLR 2021 · 被引用 122 次
- Dense-Exponential Random Features: Sharp Positive Estimators of the Gaussian KernelValerii Likhosherstov, Krzysztof Marcin Choromanski, Kumar Avinava Dubey, Frederick Liu 等NeurIPS 2023 · 被引用 5 次
- Proxyformer: Nyström-Based Linear Transformer with Trainable Proxy TokensSangho Lee, Hayun Lee, Dongkun ShinAAAI 2024 · 被引用 5 次
