PermuteFormer: Efficient Relative Position Encoding for Long Sequences
Peng Chen
Abstract
A recent variation of Transformer, Performer, scales Transformer to longer sequences with a linear attention mechanism. However, it is not compatible with relative position encoding, which has advantages over absolute position encoding. In this paper, we discuss possible ways to add relative position encoding to Performer. Based on the analysis, we propose Per-muteFormer, a Performer-based model with relative position encoding that scales linearly on long sequences. PermuteFormer applies position-dependent transformation on queries and keys to encode positional information into the attention module. This transformation is carefully crafted so that the final output of selfattention is not affected by absolute positions of tokens. PermuteFormer introduces negligible computational overhead by design that it runs as fast as Performer. We evaluate Per-muteFormer on Long-Range Arena, a dataset for long sequences, as well as WikiText-103, a language modeling dataset. The experiments show that PermuteFormer uniformly improves the performance of Performer with almost no computational overhead and outperforms vanilla Transformer on most of the tasks. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cdcb9d41-f3f6-45ad-ba71-75adbb56e2a2Cited by top-tier papers2
- Ripple Attention for Visual Perception with Sub-quadratic ComplexityLin Zheng, Huijie Pan, Lingpeng KongICML 2022 · 4 citations
- Efficient Attention via Control VariatesLin Zheng, Jianbo Yuan, Chong Wang, Lingpeng KongICLR 2023 · 2 citations
Builds on12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
Related papers
- Toeplitz Neural Network for Sequence ModelingZhen Qin, Xiaodong Han, Weixuan Sun, Bowen He et al.ICLR 2023 · 5 citations
- CAPE: Encoding Relative Positions with Continuous Augmented Positional EmbeddingsTatiana Likhomanenko, Qiantong Xu, Gabriel Synnaeve, Ronan Collobert et al.NeurIPS 2021 · 74 citations
- A Simple and Effective Positional Encoding for TransformersPu-Chin Chen, Henry Tsai, Srinadh Bhojanapalli, Hyung Won Chung et al.EMNLP 2021 · 51 citations
- Relative Positional Encoding for Transformers with Linear ComplexityAntoine Liutkus, Ondrej Cífka, Shih-Lun Wu, Umut Simsekli et al.ICML 2021 · 63 citations
- Functional Interpolation for Relative Positions improves Long Context TransformersShanda Li, Chong You, Guru Guruganesh, Joshua Ainslie et al.ICLR 2024 · 66 citations
