On Scalar Embedding of Relative Positions in Attention Models
Junshuang Wu, Richong Zhang, Yongyi Mao, Junfan Chen
摘要
Attention with positional encoding has been demonstrated as a powerful component in modern neural network models, such as transformers. However, why positional encoding works well in attention models remains largely unanswered. In this paper, we study the scalar relative positional encoding (SRPE) proposed in the T5 transformer. Such an encoding method has two features. First, it uses a scalar to embed relative positions. Second, the relative positions are bucketized using a fixed heuristic algorithm, and positions in the same bucket share the same embedding. In this work, we show that SRPE in attention has an elegant probabilistic interpretation. More specifically, the positional encoding serves to produce a prior distribution for the attended positions. The resulting attentive distribution can be viewed as a posterior distribution of the attended position given the observed input sequence. Furthermore, we propose a new SRPE (AT5) that adopts a learnable bucketization protocol and automatically adapts to the dependency range specific to the learning task. Empirical studies show that the AT5 achieves superior performance than the T5's SRPE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- Rethinking Positional Encoding in Language Pre-trainingGuolin Ke, Di He, Tie-Yan LiuICLR 2021 · 被引用 358 次
- Encoding word order in complex embeddingsBenyou Wang, Donghao Zhao, Christina Lioma, Qiuchi Li 等ICLR 2020 · 被引用 134 次
- Multiple Positional Self-Attention Network for Text ClassificationBiyun Dai, Jinlong Li, Ruoyi XuAAAI 2020 · 被引用 10 次
相关 Paper
- A Simple and Effective Positional Encoding for TransformersPu-Chin Chen, Henry Tsai, Srinadh Bhojanapalli, Hyung Won Chung 等EMNLP 2021 · 被引用 51 次
- DAPE: Data-Adaptive Positional Encoding for Length ExtrapolationChuanyang Zheng, Yihang Gao, Han Shi, Minbin Huang 等NeurIPS 2024 · 被引用 42 次
- Conditional Positional Encodings for Vision TransformersXiangxiang Chu, Zhi Tian, Bo Zhang, Xinlong Wang 等ICLR 2023 · 被引用 406 次
- The Impact of Positional Encoding on Length Generalization in TransformersAmirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das 等NeurIPS 2023 · 被引用 444 次
- Positional Encoding for Spiking TransformersZijian Zhou, Yu Liang, Honglin Cao, Ammar Belatreche 等ICML 2026 · 被引用 7 次
