HEAL-SWIN: A Vision Transformer on the Sphere
Oscar Carlsson, Jan E. Gerken, Hampus Linander, Heiner Spieß, Fredrik Ohlsson, Christoffer Petersson, Daniel Persson
Abstract
High-resolution wide-angle fisheye images are becoming more and more important for robotics applications such as autonomous driving. However, using ordinary convolutional neural networks or vision transformers on this data is problematic due to projection and distortion losses introduced when projecting to a rectangular grid on the plane. We introduce the HEAL-SWIN transformer, which combines the highly uniform Hierarchical Equal Area iso-Latitude Pixelation (HEALPix) grid used in astrophysics and cosmology with the Hierarchical Shifted-Window (SWIN) transformer to yield an efficient and flexible model capable of training on high-resolution, distortion-free spherical data. In HEAL-SWIN, the nested structure of the HEALPix grid is used to perform the patching and windowing operations of the SWIN transformer, enabling the network to process spherical representations with minimal computational overhead. We demonstrate the superior performance of our model on both synthetic and real automotive datasets, as well as a selection of other image datasets, for semantic segmentation, depth regression and classification tasks. Our code is publicly available 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 55b87a55-4b7a-443d-bff5-8cce8ea52c22Cited by top-tier papers4
- Efficiency Follows Global-Local DecouplingZhenyu Yang, Gensheng Pei, Tao Chen, Yichao Zhou et al.CVPR 2026 · 3 citations
- Spherical SO(3) Equivariant Local AttentionYusuke Sekikawa, Jun Nagata, Itsumi Araki, Ruka EtoICML 2026 · 2 citations
- SCE-Depth: A Spherical Compound Eye Framework for Wide FOV Depth EstimationYi Zhu, Hao Xiong, Lin Xiao, Ranfeng Shi et al.CVPR 2026
- SphereUFormer: A U-Shaped Transformer for Spherical 360 PerceptionYaniv Benny, Lior WolfCVPR 2025
Builds on15
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
- Voxel Transformer for 3D Object DetectionJiageng Mao, Yujing Xue, Minzhe Niu, Haoyue Bai et al.ICCV 2021 · 535 citations
- Stratified Transformer for 3D Point Cloud SegmentationXin Lai, Jianhui Liu, Li Jiang, Liwei Wang et al.CVPR 2022 · 494 citations
- ClimaX: A foundation model for weather and climateTung Nguyen, Johannes Brandstetter, Ashish Kapoor, Jayesh K. Gupta et al.ICML 2023 · 426 citations
Related papers
- DarSwin: Distortion Aware Radial Swin TransformerAkshaya Athwale, Arman Afrasiyabi, Justin Lagüe, Ichrak Shili et al.ICCV 2023 · 13 citations
- TULIP: Transformer for Upsampling of LiDAR Point CloudsBin Yang, Patrick Pfreundschuh, Roland Siegwart, Marco Hutter et al.CVPR 2024 · 18 citations
- SimFIR: A Simple Framework for Fisheye Image Rectification with Self-supervised Representation LearningHao Feng, Wendi Wang, Jiajun Deng, Wengang Zhou et al.ICCV 2023 · 28 citations
- PanoSwin: a Pano-style Swin Transformer for Panorama UnderstandingZhixin Ling, Zhen Xing, Xiangdong Zhou, Manliang Cao et al.CVPR 2023
- Fish2Mesh Transformer: 3D Human Mesh Recovery from Egocentric VisionTianma Shen, Aditya Puranik, James Vong, Vrushabh Abhijit Deogirikar et al.ICCV 2025 · 1 citation
