What Do Position Embeddings Learn? An Empirical Study of Pre-Trained Language Model Positional Encoding
Yu-An Wang, Yun-Nung Chen
Abstract
In recent years, pre-trained Transformers have dominated the majority of NLP benchmark tasks. Many variants of pre-trained Transformers have kept breaking out, and most focus on designing different pre-training objectives or variants of self-attention. Embedding the position information in the self-attention mechanism is also an indispensable factor in Transformers however is often discussed at will. Therefore, this paper carries out an empirical study on position embeddings of mainstream pre-trained Transformers, which mainly focuses on two questions: 1) Do position embeddings really learn the meaning of positions? 2) How do these different learned position embeddings affect Transformers for NLP tasks? This paper focuses on providing a new insight of pre-trained position embeddings through feature-level analysis and empirical experiments on most of iconic NLP tasks. It is believed that our experimental results can guide the future work to choose the suitable positional encoding function for specific tasks given the application property. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02ed3f33-6a53-4573-9f82-41b357e7dd56Cited by top-tier papers18
- On Position Embeddings in BERTBenyou Wang, Lifeng Shang, Christina Lioma, Xin Jiang et al.ICLR 2021 · 129 citations
- Relative Positional Encoding for Transformers with Linear ComplexityAntoine Liutkus, Ondrej Cífka, Shih-Lun Wu, Umut Simsekli et al.ICML 2021 · 63 citations
- ATTEMPT: Parameter-Efficient Multi-task Tuning via Attentional Mixtures of Soft PromptsAkari Asai, Mohammadreza Salehi, Matthew E. Peters, Hannaneh HajishirziEMNLP 2022 · 55 citations
- A Simple and Effective Positional Encoding for TransformersPu-Chin Chen, Henry Tsai, Srinadh Bhojanapalli, Hyung Won Chung et al.EMNLP 2021 · 51 citations
- HiTKG: Towards Goal-Oriented Conversations via Multi-Hierarchy LearningJinjie Ni, Vlad Pandelea, Tom Young, Haicang Zhou et al.AAAI 2022 · 35 citations
Builds on1
Related papers
- Rethinking Positional Encoding in Language Pre-trainingGuolin Ke, Di He, Tie-Yan LiuICLR 2021 · 358 citations
- Position Prediction as an Effective Pretraining StrategyShuangfei Zhai, Navdeep Jaitly, Jason Ramapuram, Dan Busbridge et al.ICML 2022 · 30 citations
- Word Order Does Matter and Shuffled Language Models Know ItMostafa Abdou, Vinit Ravishankar, Artur Kulmizev, Anders SøgaardACL 2022
- Extending the Context of Pretrained LLMs by Dropping Their Positional EmbeddingYoav Gelberg, Koshi Eguchi, Takuya Akiba, Edoardo CetinICLR 2026 · 13 citations
- Enhancing Document Understanding with Group Position Embedding: A Novel Approach to Incorporate Layout InformationYuke Zhu, Yue Zhang, Dongdong Liu, Chi Xie et al.ICLR 2025
