What Rotary Position Embedding Can Tell Us: Identifying Query and Key Weights Corresponding to Basic Syntactic or High-level Semantic Information
Yiting Chen, Junchi Yan
摘要
Transformer-based large language models (LLMs) have successfully handled various tasks. As one fundamental module in Transformers, position encoding encodes the positional information of tokens in a sequence. Specifically, rotary position embedding (RoPE), one of the most widely used techniques, encodes the positional information by dividing the query or key value with d elements into d/ 2 pairs and rotating the 2d vectors corresponding to each pair of elements. Therefore, the direction of each pair and the position-related rotation jointly determine the attention score. In this paper, we show that the direction of the 2d pair is largely affected by the angle between the corresponding weight vector pair. We theoretically show that non-orthogonal weight vector pairs lead to great attention on tokens at a certain relative position and are less sensitive to the input which may correspond to basic syntactic information. Meanwhile, the orthogonal weight vector pairs are more flexible regarding the relative position, which may correspond to high-level syntactic information. Empirical evidence supports the hypothesis that shallow layers of LLMs focus more on local syntax and deep layers focus more on high-level semantics. Furthermore, we show that LLMs fine-tuning mainly changes the pairs of weight vectors that are nearly orthogonal, i.e., the weight corresponding to high-level semantics, which enables the reduction of the number of trainable parameters during fine-tuning without sacrificing performance. We propose a method namely Angle-based Weight Masking (AWM) to reduce the fine-tuning overhead and verify the effectiveness of the proposed method on widely used Alpaca fine-tuned Llama-2.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- A Circular Argument: Does RoPE need to be Equivariant for Vision?Chase van de Geijn, Timo Lüddecke, Polina Turishcheva, Alexander S. EckerNeurIPS 2025 · 被引用 6 次
- Decoupling Positional and Symbolic Attention in TransformersFelipe Urrutia, Jorge Salas, Alexander Kozachinskiy, Cristian Buc Calderon 等ICLR 2026 · 被引用 3 次
- Predictable Compression Failures: Order Sensitivity and Information Budgeting for Evidence-Grounded Binary AdjudicationLeon Chlon, Ahmed Karim, MarcAntonio AwadaICML 2026 · 被引用 2 次
它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 被引用 508 次
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language ModelsYung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim 等ICLR 2024 · 被引用 354 次
相关 Paper
- Round and Round We Go! What makes Rotary Positional Encodings useful?Federico Barbero, Alex Vitvitskyi, Christos Perivolaropoulos, Razvan Pascanu 等ICLR 2025
- Base of RoPE Bounds Context LengthMingyu Xu, Xin Men, Bingning Wang, Qingyu Zhang 等NeurIPS 2024 · 被引用 56 次
- Massive Values in Self-Attention Modules are the Key to Contextual Knowledge UnderstandingMingyu Jin, Kai Mei, Wujiang Xu, Mingjie Sun 等ICML 2025
- Deconstructing Positional Information: From Attention Logits to Training BiasesZihan Gu, Ruoyu Chen, Han Zhang, Hua Zhang 等ICLR 2026 · 被引用 4 次
- Dynamic Positional Attention Modulation for Parameter-Efficient Fine-Tuning of Large Language ModelsDayan Pan, Jingyuan Wang, Xie YuKDD 2026
