Exploring Transformer Extrapolation
Zhen Qin, Yiran Zhong, Hui Deng
摘要
Length extrapolation has attracted considerable attention recently since it allows transformers to be tested on longer sequences than those used in training. Previous research has shown that this property can be attained by using carefully designed Relative Positional Encodings (RPEs). While these methods perform well on a variety of corpora, the conditions for length extrapolation have yet to be investigated. This paper attempts to determine what types of RPEs allow for length extrapolation through a thorough mathematical and empirical analysis. We discover that a transformer is certain to possess this property as long as the series that corresponds to the RPE's exponential converges. Two practices are derived from the conditions and examined in language modeling tasks on a variety of corpora. As a bonus from the conditions, we derive a new Theoretical Receptive Field (TRF) to measure the receptive field of RPEs without taking any training steps. Extensive experiments are conducted on the Wikitext-103, Books, Github, and WikiBook datasets to demonstrate the viability of our discovered conditions. We also compare TRF to Empirical Receptive Field (ERF) across different models, showing consistently matched trends on these datasets. Code is released at: https://github.com/OpenNLPLab/Rpe.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Titans: Learning to Memorize at Test TimeAli Behrouz, Peilin Zhong, Vahab MirrokniNeurIPS 2025 · 被引用 368 次
- DAPE: Data-Adaptive Positional Encoding for Length ExtrapolationChuanyang Zheng, Yihang Gao, Han Shi, Minbin Huang 等NeurIPS 2024 · 被引用 42 次
- Various Lengths, Constant Speed: Efficient Language Modeling with Lightning AttentionZhen Qin, Weigao Sun, Dong Li, Xuyang Shen 等ICML 2024 · 被引用 26 次
- DAPE V2: Process Attention Score as Feature Map for Length ExtrapolationChuanyang Zheng, Yihang Gao, Han Shi, Jing Xiong 等ACL 2025 · 被引用 12 次
- Navigating Scaling Laws: Compute Optimality in Adaptive Model TrainingSotiris Anagnostidis, Gregor Bachmann, Imanol Schlag, Thomas HofmannICML 2024 · 被引用 2 次
它引用的顶会 Paper7
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 被引用 1,168 次
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and TextHassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang 等NeurIPS 2021 · 被引用 782 次
- Hierarchically Gated Recurrent Neural Network for Sequence ModelingZhen Qin, Songlin Yang, Yiran ZhongNeurIPS 2023 · 被引用 152 次
- KERPLE: Kernelized Relative Positional Embedding for Length ExtrapolationTa-Chung Chi, Ting-Han Fan, Peter J. Ramadge, Alexander RudnickyNeurIPS 2022 · 被引用 112 次
- Relative Positional Encoding for Transformers with Linear ComplexityAntoine Liutkus, Ondrej Cífka, Shih-Lun Wu, Umut Simsekli 等ICML 2021 · 被引用 63 次
相关 Paper
- Dissecting Transformer Length Extrapolation via the Lens of Receptive Field AnalysisTa-Chung Chi, Ting-Han Fan, Alexander Rudnicky, Peter J. RamadgeACL 2023 · 被引用 4 次
- Functional Interpolation for Relative Positions improves Long Context TransformersShanda Li, Chong You, Guru Guruganesh, Joshua Ainslie 等ICLR 2024 · 被引用 66 次
- Your Transformer May Not be as Powerful as You ExpectShengjie Luo, Shanda Li, Shuxin Zheng, Tie-Yan Liu 等NeurIPS 2022 · 被引用 69 次
- A Length-Extrapolatable TransformerYutao Sun, Li Dong, Barun Patra, Shuming Ma 等ACL 2023 · 被引用 45 次
- Context-aware Biases for Length ExtrapolationAli Veisi, Hamidreza Amirzadeh, Amir MansourianEMNLP 2025 · 被引用 2 次
