Dissecting Transformer Length Extrapolation via the Lens of Receptive Field Analysis
Ta-Chung Chi, Ting-Han Fan, Alexander Rudnicky, Peter J. Ramadge
Abstract
Length extrapolation permits training a transformer language model on short sequences that preserves perplexities when tested on substantially longer sequences. A relative positional embedding design, ALiBi, has had the widest usage to date. We dissect ALiBi via the lens of receptive field analysis empowered by a novel cumulative normalized gradient tool. The concept of receptive field further allows us to modify the vanilla Sinusoidal positional embedding to create Sandwich, the first parameter-free relative positional embedding design that truly length information uses longer than the training sequence. Sandwich shares with KERPLE and T5 the same logarithmic decaying temporal bias pattern with learnable relative positional embeddings; these elucidate future extrapolatable positional embedding design.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b506fbcd-27ff-48e6-abc5-c343f93afa25Cited by top-tier papers23
- The Impact of Positional Encoding on Length Generalization in TransformersAmirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das et al.NeurIPS 2023 · 444 citations
- Transformers Can Do Arithmetic with the Right EmbeddingsSean McLeish, Arpit Bansal, Alex Stein, Neel Jain et al.NeurIPS 2024 · 94 citations
- Training-Free Long-Context Scaling of Large Language ModelsChenxin An, Fei Huang, Jun Zhang, Shansan Gong et al.ICML 2024 · 68 citations
- Functional Interpolation for Relative Positions improves Long Context TransformersShanda Li, Chong You, Guru Guruganesh, Joshua Ainslie et al.ICLR 2024 · 66 citations
- Base of RoPE Bounds Context LengthMingyu Xu, Xin Men, Bingning Wang, Qingyu Zhang et al.NeurIPS 2024 · 56 citations
Builds on7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang et al.NeurIPS 2021 · 1,553 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- ETC: Encoding Long and Structured Inputs in TransformersJoshua Ainslie, Santiago Ontañón, Chris Alberti, Vaclav Cvicek et al.EMNLP 2020 · 268 citations
Related papers
- Exploring Transformer ExtrapolationZhen Qin, Yiran Zhong, Hui DengAAAI 2024 · 12 citations
- Wavelet-based Positional Representation for Long ContextYui Oka, Taku Hasegawa, Kyosuke Nishida, Kuniko SaitoICLR 2025
- A Length-Extrapolatable TransformerYutao Sun, Li Dong, Barun Patra, Shuming Ma et al.ACL 2023 · 45 citations
- KERPLE: Kernelized Relative Positional Embedding for Length ExtrapolationTa-Chung Chi, Ting-Han Fan, Peter J. Ramadge, Alexander RudnickyNeurIPS 2022 · 112 citations
- CAPE: Encoding Relative Positions with Continuous Augmented Positional EmbeddingsTatiana Likhomanenko, Qiantong Xu, Gabriel Synnaeve, Ronan Collobert et al.NeurIPS 2021 · 74 citations
