Spherical SO(3) Equivariant Local Attention
Yusuke Sekikawa, Jun Nagata, Itsumi Araki, Ruka Eto
Abstract
Spherical signals provide a natural representation for omnidirectional perception and often benefit from equivariance to 3D rotations. Recent spherical vision transformers implement local self-attention on spherical grids, but most retain only partial equivariance and rely on local positional embeddings (LPEs). Such LPEs can degrade robustness to camera tilt or object reorientation and introduce additional memory and computational overhead. We propose \textit{Spherical SO(3)-Equivariant Local Attention} (SoLA), an LPE-free local attention mechanism for spherical signals. SoLA achieves full equivariance through a distance-preserving positional modulation that couples query/key features with each token’s unit direction. Specifically, the modulation lifts queries and keys using an outer-product with the 4D direction dependent vector. The induced similarity of the modulated queries and keys depends on content affinity and great-circle distance while remaining invariant to global rotations. The same formulation admits a softmax-free linear variant that computes local attention via key-value aggregation without per-query neighbor materialization. We integrate SoLA into a U-shaped spherical transformer for depth estimation and semantic segmentation, demonstrating substantially improved robustness to arbitrary 3D rotations compared to prior spherical transformers with similar computational costs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0023508a-c4b7-4a39-a041-3f76887b6245Builds on7
- LieTransformer: Equivariant Self-Attention for Lie GroupsMichael J. Hutchinson, Charline Le Lan, Sheheryar Zaidi, Emilien Dupont et al.ICML 2021 · 132 citations
- Orientation-Aware Semantic Segmentation on Icosahedron SpheresChao Zhang, Stephan Liwicki, William Smith, Roberto CipollaICCV 2019 · 90 citations
- Neural Operators with Localized Integral and Differential KernelsMiguel Liu-Schiaffini, Julius Berner, Boris Bonev, Thorsten Kurth et al.ICML 2024 · 63 citations
- Attention on the SphereBoris Bonev, Max Rietmann, Andrea Paris, Alberto Carpentieri et al.NeurIPS 2025 · 13 citations
- Scalable and Equivariant Spherical CNNs by Discrete-Continuous (DISCO) ConvolutionsJeremy Ocampo, Matthew A. Price, Jason D. McEwenICLR 2023 · 5 citations
Related papers
- SphereUFormer: A U-Shaped Transformer for Spherical 360 PerceptionYaniv Benny, Lior WolfCVPR 2025
- 3D Equivariant Pose Regression via Direct Wigner-D Harmonics PredictionJongmin Lee, Minsu ChoNeurIPS 2024 · 6 citations
- SoPE: Spherical Coordinate-Based Positional Embedding for Enhancing Spatial Perception of 3D LVLMsKoonting Yip, Qiyan Zhao, Wenhao Yu, Liangyu Yuan et al.CVPR 2026 · 3 citations
- SO(3)-Equivariant ViT-Adapter for Data-Efficient Zero-Shot Sim-to-Real Indoor Panoramic Depth EstimationZiyan He, Qiudan Zhang, Lin Ma, Xu WangCVPR 2026
- PanoVGGT: Feed-Forward 3D Reconstruction from Panoramic ImageryYijing Guo, Mengjun Chao, Luo Wang, Tianyang Zhao et al.CVPR 2026 · 11 citations
