Absolute Position Embedding Learns Sinusoid-like Waves for Attention Based on Relative Position
Yuji Yamamoto, Takuya Matsuzaki
摘要
Attention weight is a clue to interpret how a Transformer-based model makes an inference. In some attention heads, the attention focuses on the neighbors of each token. This allows the output vector of each token to depend on the surrounding tokens and contributes to make the inference context-dependent. We analyze the mechanism behind the concentration of attention on nearby tokens. We show that the phenomenon emerges as follows: (1) learned position embedding has sinusoid-like components, (2) such components are transmitted to the query and the key in the self-attention, (3) the attention head shifts the phases of the sinusoid-like components so that the attention concentrates on nearby tokens at specific relative positions. In other words, a certain type of Transformer-based model acquires the sinusoidal positional encoding to some extent on its own through Masked Language Modeling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- On Position Embeddings in BERTBenyou Wang, Lifeng Shang, Christina Lioma, Xin Jiang 等ICLR 2021 · 被引用 129 次
- What Do Position Embeddings Learn? An Empirical Study of Pre-Trained Language Model Positional EncodingYu-An Wang, Yun-Nung ChenEMNLP 2020 · 被引用 71 次
- The Geometry of Multilingual Language Model RepresentationsTyler A. Chang, Zhuowen Tu, Benjamin K. BergenEMNLP 2022 · 被引用 22 次
- The Impact of Positional Encodings on Multilingual CompressionVinit Ravishankar, Anders SøgaardEMNLP 2021 · 被引用 7 次
相关 Paper
- Attention is Not Only a Weight: Analyzing Transformers with Vector NormsGoro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro InuiEMNLP 2020 · 被引用 138 次
- What are you sinking? A geometric approach on attention sinkValeria Ruscio, Umberto Nanni, Fabrizio SilvestriNeurIPS 2025 · 被引用 29 次
- On Scalar Embedding of Relative Positions in Attention ModelsJunshuang Wu, Richong Zhang, Yongyi Mao, Junfan ChenAAAI 2021 · 被引用 4 次
- Scan and Snap: Understanding Training Dynamics and Token Composition in 1-layer TransformerYuandong Tian, Yiping Wang, Beidi Chen, Simon S. DuNeurIPS 2023 · 被引用 125 次
- On the Emergence of Position Bias in TransformersXinyi Wu, Yifei Wang, Stefanie Jegelka, Ali JadbabaieICML 2025 · 被引用 1 次
