Beyond Position: the emergence of wavelet-like properties in Transformers
Valeria Ruscio, Umberto Nanni, Fabrizio Silvestri
摘要
This paper studies how Transformer models with Rotary Position Embeddings (RoPE) develop emergent, wavelet-like properties that compensate for the positional encoding's theoretical limitations. Through an analysis spanning model scales, architectures, and training checkpoints, we show that attention heads evolve to implement multi-resolution processing analogous to wavelet transforms. We demonstrate that this scale-invariant behavior is unique to RoPE, emerges through distinct evolutionary phases during training, and statistically adheres to the fundamental uncertainty principle. Our findings suggest that the effectiveness of modern Transformers stems from their remarkable ability to spontaneously develop optimal, multi-resolution decompositions to address inherent architectural constraints.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- What are you sinking? A geometric approach on attention sinkValeria Ruscio, Umberto Nanni, Fabrizio SilvestriNeurIPS 2025 · 被引用 29 次
- Resonating with RoPE: Spectral Quantization for High-Fidelity Key Cache CompressionXuefei Wang, Haoyu Tang, Tianyuan Liang, Zhibin Wang 等ACL 2026
它引用的顶会 Paper2
相关 Paper
- Wavelet-based Positional Representation for Long ContextYui Oka, Taku Hasegawa, Kyosuke Nishida, Kuniko SaitoICLR 2025
- AdaRoPE: Not All Attention Heads Should Rotate and Scale EquallyShaowen Wang, Yuke Zheng, Tansheng Zhu, Shuang Chen 等ICML 2026
- Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D PlaneHaoyu Liu, Sucheng Ren, Tingyu Zhu, Peng Wang 等ICML 2026
- nD-RoPE: A Generalized RoPE for n-Dimensional Position EmbeddingBoyang Li, Yulin Wu, Sizhe Xu, Nuoxian Huang 等ICML 2026
- Functional Equivalence in Attention: A Comprehensive Study with Applications to Linear Mode ConnectivityViet Hoang Tran, VINH KHANH BUI, Van-Hoan Trinh, Ngoc Tan Lai 等ICML 2026
