Beyond Position: the emergence of wavelet-like properties in Transformers
Valeria Ruscio, Umberto Nanni, Fabrizio Silvestri
Abstract
This paper studies how Transformer models with Rotary Position Embeddings (RoPE) develop emergent, wavelet-like properties that compensate for the positional encoding's theoretical limitations. Through an analysis spanning model scales, architectures, and training checkpoints, we show that attention heads evolve to implement multi-resolution processing analogous to wavelet transforms. We demonstrate that this scale-invariant behavior is unique to RoPE, emerges through distinct evolutionary phases during training, and statistically adheres to the fundamental uncertainty principle. Our findings suggest that the effectiveness of modern Transformers stems from their remarkable ability to spontaneously develop optimal, multi-resolution decompositions to address inherent architectural constraints.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aada859f-ea8e-4a32-ae01-7746a1438138Cited by top-tier papers2
- What are you sinking? A geometric approach on attention sinkValeria Ruscio, Umberto Nanni, Fabrizio SilvestriNeurIPS 2025 · 29 citations
- Resonating with RoPE: Spectral Quantization for High-Fidelity Key Cache CompressionXuefei Wang, Haoyu Tang, Tianyuan Liang, Zhibin Wang et al.ACL 2026
Builds on2
Related papers
- Wavelet-based Positional Representation for Long ContextYui Oka, Taku Hasegawa, Kyosuke Nishida, Kuniko SaitoICLR 2025
- AdaRoPE: Not All Attention Heads Should Rotate and Scale EquallyShaowen Wang, Yuke Zheng, Tansheng Zhu, Shuang Chen et al.ICML 2026
- Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D PlaneHaoyu Liu, Sucheng Ren, Tingyu Zhu, Peng Wang et al.ICML 2026
- nD-RoPE: A Generalized RoPE for n-Dimensional Position EmbeddingBoyang Li, Yulin Wu, Sizhe Xu, Nuoxian Huang et al.ICML 2026
- Functional Equivalence in Attention: A Comprehensive Study with Applications to Linear Mode ConnectivityViet Hoang Tran, VINH KHANH BUI, Van-Hoan Trinh, Ngoc Tan Lai et al.ICML 2026
