Attention on the Sphere
Boris Bonev, Max Rietmann, Andrea Paris, Alberto Carpentieri, Thorsten Kurth
Abstract
We introduce a generalized attention mechanism for spherical domains, enabling Transformer architectures to natively process data defined on the two-dimensional sphere -a critical need in fields such as atmospheric physics, cosmology, and robotics, where preserving spherical symmetries and topology is essential for physical accuracy. By integrating numerical quadrature weights into the attention mechanism, we obtain a geometrically faithful spherical attention that is approximately rotationally equivariant, providing strong inductive biases and leading to better performance than Cartesian approaches. To further enhance both scalability and model performance, we propose neighborhood attention on the sphere, which confines interactions to geodesic neighborhoods. This approach reduces computational complexity and introduces the additional inductive bias for locality, while retaining the symmetry properties of our method. We provide optimized CUDA kernels and memory-efficient implementations to ensure practical applicability.
The method is validated on three diverse tasks: simulating shallow water equations on the rotating sphere, spherical image segmentation, and spherical depth estimation. Across all tasks, our spherical Transformers consistently outperform their planar counterparts, highlighting the advantage of geometric priors for learning on spherical domains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 74bcd171-b570-4706-b40d-a062095956d6Cited by top-tier papers1
Ask how each one uses itBuilds on15
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- SE(3)-Transformers: 3D Roto-Translation Equivariant Attention NetworksFabian Fuchs, Daniel E. Worrall, Volker Fischer, Max WellingNeurIPS 2020 · 1,025 citations
- Linear Transformers Are Secretly Fast Weight ProgrammersImanol Schlag, Kazuki Irie, Jürgen SchmidhuberICML 2021 · 394 citations
- Spherical Fourier Neural Operators: Learning Stable Dynamics on the SphereBoris Bonev, Thorsten Kurth, Christian Hundt, Jaideep Pathak et al.ICML 2023 · 280 citations
- Poseidon: Efficient Foundation Models for PDEsMaximilian Herde, Bogdan Raonic, Tobias Rohner, Roger Käppeli et al.NeurIPS 2024 · 235 citations
Related papers
- SphereUFormer: A U-Shaped Transformer for Spherical 360 PerceptionYaniv Benny, Lior WolfCVPR 2025
- QUEST: A robust attention formulation using query-modulated spherical attentionHariprasath Govindarajan, Per Sidén, Jacob Roll, Fredrik LindstenICLR 2026 · 1 citation
- Spatial Attention Kinetic Networks with E(n)-EquivarianceYuanqing Wang, John D. ChoderaICLR 2023 · 11 citations
- Bridging Equivariant GNNs and Spherical CNNs for Structured Physical DomainsColin Kohler, Purvik Patel, Nathan Vaska, Justin A. Goodwin et al.NeurIPS 2025 · 1 citation
- Scaling Spherical CNNsCarlos Esteves, Jean-Jacques E. Slotine, Ameesh MakadiaICML 2023 · 28 citations
