KERPLE: Kernelized Relative Positional Embedding for Length Extrapolation
Ta-Chung Chi, Ting-Han Fan, Peter J. Ramadge, Alexander Rudnicky
摘要
Relative positional embeddings (RPE) have received considerable attention since RPEs effectively model the relative distance among tokens and enable length extrapolation. We propose KERPLE, a framework that generalizes relative position embedding for extrapolation by kernelizing positional differences. We achieve this goal using conditionally positive definite (CPD) kernels, a class of functions known for generalizing distance metrics. To maintain the inner product interpretation of self-attention, we show that a CPD kernel can be transformed into a PD kernel by adding a constant offset. This offset is implicitly absorbed in the Softmax normalization during self-attention. The diversity of CPD kernels allows us to derive various RPEs that enable length extrapolation in a principled way. Experiments demonstrate that the logarithmic variant achieves excellent extrapolation performance on three large language modeling datasets. Our implementation and pretrained checkpoints are released at https://github.com/chijames/KERPLE.git . * Equal contribution 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper42
- Transformers Can Do Arithmetic with the Right EmbeddingsSean McLeish, Arpit Bansal, Alex Stein, Neel Jain 等NeurIPS 2024 · 被引用 94 次
- Functional Interpolation for Relative Positions improves Long Context TransformersShanda Li, Chong You, Guru Guruganesh, Joshua Ainslie 等ICLR 2024 · 被引用 66 次
- Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In ContextXiang Cheng, Yuxin Chen, Suvrit SraICML 2024 · 被引用 64 次
- Rope to Nope and Back Again: A New Hybrid Attention StrategyBowen Yang, Bharat Venkitesh, Dwaraknath Gnaneshwar, Hangyu Lin 等NeurIPS 2025 · 被引用 51 次
- DAPE: Data-Adaptive Positional Encoding for Length ExtrapolationChuanyang Zheng, Yihang Gao, Han Shi, Minbin Huang 等NeurIPS 2024 · 被引用 42 次
它引用的顶会 Paper15
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 被引用 1,168 次
相关 Paper
- Context-aware Biases for Length ExtrapolationAli Veisi, Hamidreza Amirzadeh, Amir MansourianEMNLP 2025 · 被引用 2 次
- Exploring Transformer ExtrapolationZhen Qin, Yiran Zhong, Hui DengAAAI 2024 · 被引用 12 次
- Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length ExtrapolationZhenyu He, Guhao Feng, Shengjie Luo, Kai Yang 等ICML 2024 · 被引用 25 次
- Dissecting Transformer Length Extrapolation via the Lens of Receptive Field AnalysisTa-Chung Chi, Ting-Han Fan, Alexander Rudnicky, Peter J. RamadgeACL 2023 · 被引用 4 次
- A Length-Extrapolatable TransformerYutao Sun, Li Dong, Barun Patra, Shuming Ma 等ACL 2023 · 被引用 45 次
