PermuteFormer: Efficient Relative Position Encoding for Long Sequences
Peng Chen
摘要
A recent variation of Transformer, Performer, scales Transformer to longer sequences with a linear attention mechanism. However, it is not compatible with relative position encoding, which has advantages over absolute position encoding. In this paper, we discuss possible ways to add relative position encoding to Performer. Based on the analysis, we propose Per-muteFormer, a Performer-based model with relative position encoding that scales linearly on long sequences. PermuteFormer applies position-dependent transformation on queries and keys to encode positional information into the attention module. This transformation is carefully crafted so that the final output of selfattention is not affected by absolute positions of tokens. PermuteFormer introduces negligible computational overhead by design that it runs as fast as Performer. We evaluate Per-muteFormer on Long-Range Arena, a dataset for long sequences, as well as WikiText-103, a language modeling dataset. The experiments show that PermuteFormer uniformly improves the performance of Performer with almost no computational overhead and outperforms vanilla Transformer on most of the tasks. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Ripple Attention for Visual Perception with Sub-quadratic ComplexityLin Zheng, Huijie Pan, Lingpeng KongICML 2022 · 被引用 4 次
- Efficient Attention via Control VariatesLin Zheng, Jianbo Yuan, Chong Wang, Lingpeng KongICLR 2023 · 被引用 2 次
它引用的顶会 Paper12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
相关 Paper
- Toeplitz Neural Network for Sequence ModelingZhen Qin, Xiaodong Han, Weixuan Sun, Bowen He 等ICLR 2023 · 被引用 5 次
- CAPE: Encoding Relative Positions with Continuous Augmented Positional EmbeddingsTatiana Likhomanenko, Qiantong Xu, Gabriel Synnaeve, Ronan Collobert 等NeurIPS 2021 · 被引用 74 次
- A Simple and Effective Positional Encoding for TransformersPu-Chin Chen, Henry Tsai, Srinadh Bhojanapalli, Hyung Won Chung 等EMNLP 2021 · 被引用 51 次
- Relative Positional Encoding for Transformers with Linear ComplexityAntoine Liutkus, Ondrej Cífka, Shih-Lun Wu, Umut Simsekli 等ICML 2021 · 被引用 63 次
- Functional Interpolation for Relative Positions improves Long Context TransformersShanda Li, Chong You, Guru Guruganesh, Joshua Ainslie 等ICLR 2024 · 被引用 66 次
