Computation Mechanism Behind LLM Position Generalization
Chi Han, Heng Ji
Abstract
Most written natural languages are composed of sequences of words and sentences. Similar to humans, large language models (LLMs) exhibit flexibility in handling textual positions -a phenomenon we term position generalization. They can understand texts with position perturbations and generalize to longer texts than those encountered during training with the latest techniques. These phenomena suggest that LLMs handle positions tolerantly, but how LLMs computationally process positional relevance remains largely unexplored. This work connects the linguistic phenomenon with LLMs' computational mechanisms. We show how LLMs enforce certain computational mechanisms for the aforementioned tolerance in position perturbations. Despite the complex design of the self-attention mechanism, this work reveals that LLMs learn a counterintuitive disentanglement of attention logits. Their values show a 0.959 linear correlation with an approximation of the arithmetic sum of positional relevance and semantic importance. Furthermore, we identify a prevalent pattern in intermediate features, which we prove theoretically enables this effect. The pattern, which is different from how randomly initialized parameters would behave, suggests that it is a learned behavior rather than a natural result of the model architecture. Based on these findings, we provide computational explanations and criteria for LLMs' position flexibilities. This work takes a pioneering step in linking position generalization with modern LLMs' internal mechanism.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4c202d01-c987-40d5-ae37-17312302d17dCited by top-tier papers1
Ask how each one uses itBuilds on4
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context MemoryChaojun Xiao, Pengle Zhang, Xu Han, Guangxuan Xiao et al.NeurIPS 2024 · 223 citations
- In-Context Pretraining: Language Modeling Beyond Document BoundariesWeijia Shi, Sewon Min, Maria Lomeli, Chunting Zhou et al.ICLR 2024 · 87 citations
- Eliminating Position Bias of Language Models: A Mechanistic ApproachZiqi Wang, Hanlin Zhang, Xiner Li, Kuan-Hao Huang et al.ICLR 2025
Related papers
- Exploring Context Window of Large Language Models via Decomposed Positional VectorsZican Dong, Junyi Li, Xin Men, Xin Zhao et al.NeurIPS 2024 · 37 citations
- A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit InterpolationYan Li, Tianyi Zhang, Zechuan Li, Caren HanEMNLP 2025
- The Role of Sparsity for Length Generalization in LLMsNoah Golowich, Samy Jelassi, David Brandfonbrener, Sham M. Kakade et al.ICML 2025
- Long-Short Alignment for Effective Long-Context Modeling in LLMsTianqi Du, Haotian Huang, Yifei Wang, Yisen WangICML 2025
- Word Order Does Matter and Shuffled Language Models Know ItMostafa Abdou, Vinit Ravishankar, Artur Kulmizev, Anders SøgaardACL 2022
