Spherical SO(3) Equivariant Local Attention
Yusuke Sekikawa, Jun Nagata, Itsumi Araki, Ruka Eto
摘要
Spherical signals provide a natural representation for omnidirectional perception and often benefit from equivariance to 3D rotations. Recent spherical vision transformers implement local self-attention on spherical grids, but most retain only partial equivariance and rely on local positional embeddings (LPEs). Such LPEs can degrade robustness to camera tilt or object reorientation and introduce additional memory and computational overhead. We propose \textit{Spherical SO(3)-Equivariant Local Attention} (SoLA), an LPE-free local attention mechanism for spherical signals. SoLA achieves full equivariance through a distance-preserving positional modulation that couples query/key features with each token’s unit direction. Specifically, the modulation lifts queries and keys using an outer-product with the 4D direction dependent vector. The induced similarity of the modulated queries and keys depends on content affinity and great-circle distance while remaining invariant to global rotations. The same formulation admits a softmax-free linear variant that computes local attention via key-value aggregation without per-query neighbor materialization. We integrate SoLA into a U-shaped spherical transformer for depth estimation and semantic segmentation, demonstrating substantially improved robustness to arbitrary 3D rotations compared to prior spherical transformers with similar computational costs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- LieTransformer: Equivariant Self-Attention for Lie GroupsMichael J. Hutchinson, Charline Le Lan, Sheheryar Zaidi, Emilien Dupont 等ICML 2021 · 被引用 132 次
- Orientation-Aware Semantic Segmentation on Icosahedron SpheresChao Zhang, Stephan Liwicki, William Smith, Roberto CipollaICCV 2019 · 被引用 90 次
- Neural Operators with Localized Integral and Differential KernelsMiguel Liu-Schiaffini, Julius Berner, Boris Bonev, Thorsten Kurth 等ICML 2024 · 被引用 63 次
- Attention on the SphereBoris Bonev, Max Rietmann, Andrea Paris, Alberto Carpentieri 等NeurIPS 2025 · 被引用 13 次
- Scalable and Equivariant Spherical CNNs by Discrete-Continuous (DISCO) ConvolutionsJeremy Ocampo, Matthew A. Price, Jason D. McEwenICLR 2023 · 被引用 5 次
相关 Paper
- SphereUFormer: A U-Shaped Transformer for Spherical 360 PerceptionYaniv Benny, Lior WolfCVPR 2025
- 3D Equivariant Pose Regression via Direct Wigner-D Harmonics PredictionJongmin Lee, Minsu ChoNeurIPS 2024 · 被引用 6 次
- SoPE: Spherical Coordinate-Based Positional Embedding for Enhancing Spatial Perception of 3D LVLMsKoonting Yip, Qiyan Zhao, Wenhao Yu, Liangyu Yuan 等CVPR 2026 · 被引用 3 次
- SO(3)-Equivariant ViT-Adapter for Data-Efficient Zero-Shot Sim-to-Real Indoor Panoramic Depth EstimationZiyan He, Qiudan Zhang, Lin Ma, Xu WangCVPR 2026
- PanoVGGT: Feed-Forward 3D Reconstruction from Panoramic ImageryYijing Guo, Mengjun Chao, Luo Wang, Tianyang Zhao 等CVPR 2026 · 被引用 11 次
