On Scalar Embedding of Relative Positions in Attention Models
Junshuang Wu, Richong Zhang, Yongyi Mao, Junfan Chen
Abstract
Attention with positional encoding has been demonstrated as a powerful component in modern neural network models, such as transformers. However, why positional encoding works well in attention models remains largely unanswered. In this paper, we study the scalar relative positional encoding (SRPE) proposed in the T5 transformer. Such an encoding method has two features. First, it uses a scalar to embed relative positions. Second, the relative positions are bucketized using a fixed heuristic algorithm, and positions in the same bucket share the same embedding. In this work, we show that SRPE in attention has an elegant probabilistic interpretation. More specifically, the positional encoding serves to produce a prior distribution for the attended positions. The resulting attentive distribution can be viewed as a posterior distribution of the attended position given the observed input sequence. Furthermore, we propose a new SRPE (AT5) that adopts a learnable bucketization protocol and automatically adapts to the dependency range specific to the learning task. Empirical studies show that the AT5 achieves superior performance than the T5's SRPE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 11b00e95-a71a-48b8-b078-da034f3dd128Cited by top-tier papers1
Ask how each one uses itBuilds on3
- Rethinking Positional Encoding in Language Pre-trainingGuolin Ke, Di He, Tie-Yan LiuICLR 2021 · 358 citations
- Encoding word order in complex embeddingsBenyou Wang, Donghao Zhao, Christina Lioma, Qiuchi Li et al.ICLR 2020 · 134 citations
- Multiple Positional Self-Attention Network for Text ClassificationBiyun Dai, Jinlong Li, Ruoyi XuAAAI 2020 · 10 citations
Related papers
- A Simple and Effective Positional Encoding for TransformersPu-Chin Chen, Henry Tsai, Srinadh Bhojanapalli, Hyung Won Chung et al.EMNLP 2021 · 51 citations
- DAPE: Data-Adaptive Positional Encoding for Length ExtrapolationChuanyang Zheng, Yihang Gao, Han Shi, Minbin Huang et al.NeurIPS 2024 · 42 citations
- Conditional Positional Encodings for Vision TransformersXiangxiang Chu, Zhi Tian, Bo Zhang, Xinlong Wang et al.ICLR 2023 · 406 citations
- The Impact of Positional Encoding on Length Generalization in TransformersAmirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das et al.NeurIPS 2023 · 444 citations
- Positional Encoding for Spiking TransformersZijian Zhou, Yu Liang, Honglin Cao, Ammar Belatreche et al.ICML 2026 · 7 citations
