Compass-RoPE: Isotropic Rotary Position Embeddings for Vision Transformers
Chengxi Min, Wei Wang, Yao Zhao
Abstract
Recent works introduce Rotary Position Embeddings (RoPE) into vision transformers (ViTs) to enhance their extrapolation capability, i.e., maintaining performance when inference is conducted on higher resolution images. RoPE encodes positions via rotating phases whose change is controlled by frequency components. Strandard 2D RoPE does not generalize well to input resolution changes as it only applies axial frequencies separately along each individual axis. To solve this issue, Mix-RoPE combines xy‑axis frequencies, such that it can model position relations in diagonal direction. However, in practice, we observe that the learned 2D frequencies become anisotropic in their direction distributions due to the axial spectral bias in image features, limiting the extrapolation ability of ViTs. Motivated by this observation, we propose Compass‑RoPE. We replace the xy cartesian coordinates with a polar parameterization that explicitly decouples frequency scale and angle. By initializing the angle vectors uniformly over [0,2π), it ensures the isotropic direction coverage. Besides, we further introduce discrete Fourier transform (DFT) mixing for the angle vectors, allowing each transformed individual angle vector element to nest multipule angles and thus to enrich angular expressiveness. Extensive experiments on multi-resolution classification and dense prediction tasks show that our Compass-RoPE achieves more stable extrapolation performance under large-scale resolution changes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 167a7205-e858-4f1e-bde7-0056be4ae68cBuilds on20
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
Related papers
- Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D PlaneHaoyu Liu, Sucheng Ren, Tingyu Zhu, Peng Wang et al.ICML 2026
- nD-RoPE: A Generalized RoPE for n-Dimensional Position EmbeddingBoyang Li, Yulin Wu, Sizhe Xu, Nuoxian Huang et al.ICML 2026
- Decoupling The "What" and "Where" With Polar Coordinate Positional EmbeddingAnand Gopalakrishnan, Róbert Csordás, Jürgen Schmidhuber, Michael MozerICML 2026 · 7 citations
- Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image GenerationJiaye Li, Baoyou Chen, Hui Li, Zilong Dong et al.CVPR 2026
- AdaRoPE: Not All Attention Heads Should Rotate and Scale EquallyShaowen Wang, Yuke Zheng, Tansheng Zhu, Shuang Chen et al.ICML 2026
