HEAL-SWIN: A Vision Transformer on the Sphere
Oscar Carlsson, Jan E. Gerken, Hampus Linander, Heiner Spieß, Fredrik Ohlsson, Christoffer Petersson, Daniel Persson
摘要
High-resolution wide-angle fisheye images are becoming more and more important for robotics applications such as autonomous driving. However, using ordinary convolutional neural networks or vision transformers on this data is problematic due to projection and distortion losses introduced when projecting to a rectangular grid on the plane. We introduce the HEAL-SWIN transformer, which combines the highly uniform Hierarchical Equal Area iso-Latitude Pixelation (HEALPix) grid used in astrophysics and cosmology with the Hierarchical Shifted-Window (SWIN) transformer to yield an efficient and flexible model capable of training on high-resolution, distortion-free spherical data. In HEAL-SWIN, the nested structure of the HEALPix grid is used to perform the patching and windowing operations of the SWIN transformer, enabling the network to process spherical representations with minimal computational overhead. We demonstrate the superior performance of our model on both synthetic and real automotive datasets, as well as a selection of other image datasets, for semantic segmentation, depth regression and classification tasks. Our code is publicly available 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Efficiency Follows Global-Local DecouplingZhenyu Yang, Gensheng Pei, Tao Chen, Yichao Zhou 等CVPR 2026 · 被引用 3 次
- Spherical SO(3) Equivariant Local AttentionYusuke Sekikawa, Jun Nagata, Itsumi Araki, Ruka EtoICML 2026 · 被引用 2 次
- SCE-Depth: A Spherical Compound Eye Framework for Wide FOV Depth EstimationYi Zhu, Hao Xiong, Lin Xiao, Ranfeng Shi 等CVPR 2026
- SphereUFormer: A U-Shaped Transformer for Spherical 360 PerceptionYaniv Benny, Lior WolfCVPR 2025
它引用的顶会 Paper15
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao 等CVPR 2022 · 被引用 2,138 次
- Voxel Transformer for 3D Object DetectionJiageng Mao, Yujing Xue, Minzhe Niu, Haoyue Bai 等ICCV 2021 · 被引用 535 次
- Stratified Transformer for 3D Point Cloud SegmentationXin Lai, Jianhui Liu, Li Jiang, Liwei Wang 等CVPR 2022 · 被引用 494 次
- ClimaX: A foundation model for weather and climateTung Nguyen, Johannes Brandstetter, Ashish Kapoor, Jayesh K. Gupta 等ICML 2023 · 被引用 426 次
相关 Paper
- DarSwin: Distortion Aware Radial Swin TransformerAkshaya Athwale, Arman Afrasiyabi, Justin Lagüe, Ichrak Shili 等ICCV 2023 · 被引用 13 次
- TULIP: Transformer for Upsampling of LiDAR Point CloudsBin Yang, Patrick Pfreundschuh, Roland Siegwart, Marco Hutter 等CVPR 2024 · 被引用 18 次
- SimFIR: A Simple Framework for Fisheye Image Rectification with Self-supervised Representation LearningHao Feng, Wendi Wang, Jiajun Deng, Wengang Zhou 等ICCV 2023 · 被引用 28 次
- PanoSwin: a Pano-style Swin Transformer for Panorama UnderstandingZhixin Ling, Zhen Xing, Xiangdong Zhou, Manliang Cao 等CVPR 2023
- Fish2Mesh Transformer: 3D Human Mesh Recovery from Egocentric VisionTianma Shen, Aditya Puranik, James Vong, Vrushabh Abhijit Deogirikar 等ICCV 2025 · 被引用 1 次
