Functional Equivalence in Attention: A Comprehensive Study with Applications to Linear Mode Connectivity
Viet Hoang Tran, VINH KHANH BUI, Van-Hoan Trinh, Ngoc Tan Lai, Tan Nguyen
摘要
Neural network parameter spaces are inherently non-injective, as distinct parameter configurations can realize identical functions through functional equivalence. While this symmetry is well understood in classical fully connected and convolutional models, it becomes substantially more intricate in modern attention-based architectures. Existing analyses of multihead attention have largely focused on the vanilla formulation, overlooking positional encodings that fundamentally reshape architectural symmetries. In this work, we provide a formal study of functional equivalence in Transformers with positional encodings. Focusing on the two most widely used variants--sinusoidal and rotary positional encodings (RoPE)--we show that sinusoidal encodings preserve the equivalence structure of vanilla attention, whereas rotary encodings significantly reduce the symmetry group, thereby enhancing expressivity. This offers a principled explanation for the growing prominence of RoPE in practice. We further examine how positional encodings affect linear mode connectivity, and through an alignment algorithm, empirically demonstrate that the presence and variability of connectivity across Transformer settings crucially depend on the positional encoding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper29
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 被引用 1,168 次
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 被引用 750 次
相关 Paper
- Beyond Position: the emergence of wavelet-like properties in TransformersValeria Ruscio, Umberto Nanni, Fabrizio SilvestriACL 2025 · 被引用 1 次
- Deconstructing Positional Information: From Attention Logits to Training BiasesZihan Gu, Ruoyu Chen, Han Zhang, Hua Zhang 等ICLR 2026 · 被引用 4 次
- Decoupling Positional and Symbolic Attention in TransformersFelipe Urrutia, Jorge Salas, Alexander Kozachinskiy, Cristian Buc Calderon 等ICLR 2026 · 被引用 3 次
- HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and ExtrapolationYuhan Chen, Ang Lv, Jian Luan, Bin Wang 等ACL 2025
- A Circular Argument: Does RoPE need to be Equivariant for Vision?Chase van de Geijn, Timo Lüddecke, Polina Turishcheva, Alexander S. EckerNeurIPS 2025 · 被引用 6 次
