SphereUFormer: A U-Shaped Transformer for Spherical 360 Perception
Yaniv Benny, Lior Wolf
Abstract
This paper proposes a novel method for omnidirectional 360 • perception. Most common previous methods relied on equirectangular projection. This representation is easily applicable to 2D operation layers but introduces distortions into the image. Other methods attempted to remove the distortions by maintaining a sphere representation but relied on complicated convolution kernels that failed to show competitive results. In this work, we introduce a transformer-based architecture that, by incorporating a novel "Spherical Local Self-Attention" and other spherically-oriented modules, successfully operates in the spherical domain and outperforms the state-of-the-art in 360 • perception benchmarks for depth estimation and semantic segmentation. Our code is available at https: //github.com/yanivbenny/sphere_uformer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6a47bb7d-4a00-403e-ac9c-77e426e51b8dCited by top-tier papers4
- Depth Any Panoramas: A Foundation Model for Panoramic Depth EstimationXin Lin, Meixi Song, Dizhe Zhang, Wenxuan Lu et al.CVPR 2026 · 27 citations
- Attention on the SphereBoris Bonev, Max Rietmann, Andrea Paris, Alberto Carpentieri et al.NeurIPS 2025 · 13 citations
- PVDepth: Panoramic Video Depth Estimation via Geometry-Aware Spatiotemporal AdaptationChuanxin Song, Peixi PengICML 2026
- SCE-Depth: A Spherical Compound Eye Framework for Wide FOV Depth EstimationYi Zhu, Hao Xiong, Lin Xiao, Ranfeng Shi et al.CVPR 2026
Builds on17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- Uformer: A General U-Shaped Transformer for Image RestorationZhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou et al.CVPR 2022 · 1,970 citations
Related papers
- Spherical SO(3) Equivariant Local AttentionYusuke Sekikawa, Jun Nagata, Itsumi Araki, Ruka EtoICML 2026 · 2 citations
- OmniFusion: 360 Monocular Depth Estimation via Geometry-Aware FusionYuyan Li, Yuliang Guo, Zhixin Yan, Xinyu Huang et al.CVPR 2022 · 79 citations
- EGformer: Equirectangular Geometry-biased Transformer for 360 Depth EstimationIlwi Yun, Chanyong Shin, Hyunku Lee, Hyuk-Jae Lee et al.ICCV 2023 · 50 citations
- Improving 360 Monocular Depth Estimation via Non-local Dense Prediction Transformer and Joint Supervised and Self-Supervised LearningIlwi Yun, Hyuk-Jae Lee, Chae-Eun RheeAAAI 2022 · 34 citations
- Elite360D: Towards Efficient 360 Depth Estimation via Semantic- and Distance-Aware Bi-Projection FusionHao Ai, Lin WangCVPR 2024
