Relative Positional Encoding for Transformers with Linear Complexity
Antoine Liutkus, Ondrej Cífka, Shih-Lun Wu, Umut Simsekli, Yi-Hsuan Yang, Gaël Richard
摘要
Recent advances in Transformer models allow for unprecedented sequence lengths, due to linear space and time complexity. In the meantime, relative positional encoding (RPE) was proposed as beneficial for classical Transformers and consists in exploiting lags instead of absolute positions for inference. Still, RPE is not available for the recent linear-variants of the Transformer, because it requires the explicit computation of the attention matrix, which is precisely what is avoided by such methods. In this paper, we bridge this gap and present Stochastic Positional Encoding as a way to generate PE that can be used as a replacement to the classical additive (sinusoidal) PE and provably behaves like RPE. The main theoretical contribution is to make a connection between positional encoding and cross-covariance structures of correlated Gaussian processes. We illustrate the performance of our approach on the Long-Range Arena benchmark and on music generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- cosFormer: Rethinking Softmax In AttentionZhen Qin, Weixuan Sun, Hui Deng, Dongxu Li 等ICLR 2022 · 被引用 303 次
- Learnable Fourier Features for Multi-dimensional Spatial Positional EncodingYang Li, Si Si, Gang Li, Cho-Jui Hsieh 等NeurIPS 2021 · 被引用 171 次
- A Text Attention Network for Spatial Deformation Robust Scene Text Image Super-resolutionJianqi Ma, Zhetong Liang, Lei ZhangCVPR 2022 · 被引用 95 次
- SAPE: Spatially-Adaptive Progressive Encoding for Neural OptimizationAmir Hertz, Or Perel, Raja Giryes, Olga Sorkine-Hornung 等NeurIPS 2021 · 被引用 86 次
- Functional Interpolation for Relative Positions improves Long Context TransformersShanda Li, Chong You, Guru Guruganesh, Joshua Ainslie 等ICLR 2024 · 被引用 66 次
它引用的顶会 Paper13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen 等ICLR 2021 · 被引用 881 次
相关 Paper
- PermuteFormer: Efficient Relative Position Encoding for Long SequencesPeng ChenEMNLP 2021 · 被引用 16 次
- Toeplitz Neural Network for Sequence ModelingZhen Qin, Xiaodong Han, Weixuan Sun, Bowen He 等ICLR 2023 · 被引用 5 次
- Decoupling The "What" and "Where" With Polar Coordinate Positional EmbeddingAnand Gopalakrishnan, Róbert Csordás, Jürgen Schmidhuber, Michael MozerICML 2026 · 被引用 7 次
- Stable, Fast and Accurate: Kernelized Attention with Relative Positional EncodingShengjie Luo, Shanda Li, Tianle Cai, Di He 等NeurIPS 2021 · 被引用 66 次
- Your Transformer May Not be as Powerful as You ExpectShengjie Luo, Shanda Li, Shuxin Zheng, Tie-Yan Liu 等NeurIPS 2022 · 被引用 69 次
