OmniFusion: 360 Monocular Depth Estimation via Geometry-Aware Fusion
Yuyan Li, Yuliang Guo, Zhixin Yan, Xinyu Huang, Ye Duan, Liu Ren
Abstract
A well-known challenge in applying deep-learning methods to omnidirectional images is spherical distortion. In dense regression tasks such as depth estimation, where structural details are required, using a vanilla CNN layer on the distorted 360 image results in undesired information loss. In this paper, we propose a 360 monocular depth estimation pipeline, OmniFusion, to tackle the spherical distortion issue. Our pipeline transforms a 360 image into less-distorted perspective patches (i.e. tangent images) to obtain patch-wise predictions via CNN, and then merge the patch-wise results for final output. To handle the discrepancy between patch-wise predictions which is a major issue affecting the merging quality, we propose a new framework with the following key components. First, we propose a geometry-aware feature fusion mechanism that combines 3D geometric features with 2D image features to compensate for the patch-wise discrepancy. Second, we employ the self-attention-based transformer architecture to conduct a global aggregation of patch-wise information, which further improves the consistency. Last, we introduce an iterative depth refinement mechanism, to further refine the estimated depth based on the more accurate geometric features. Experiments show that our method greatly mitigates the distortion issue, and achieves state-of-the-art performances on several 360 monocular depth estimation benchmark datasets. Our code is available at https: //github.com/yuyanli0831/OmniFusion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers26
- 360MonoDepth: High-Resolution 360° Monocular Depth EstimationManuel Rey-Area, Mingze Yuan, Christian RichardtCVPR 2022 · 80 citations
- Depth Anywhere: Enhancing 360 Monocular Depth Estimation via Perspective Distillation and Unlabeled Data AugmentationNing-Hsu Wang, Yu-Lun LiuNeurIPS 2024 · 56 citations
- PanoGRF: Generalizable Spherical Radiance Fields for Wide-baseline PanoramasZheng Chen, Yan-Pei Cao, Yuan-Chen Guo, Chen Wang et al.NeurIPS 2023 · 28 citations
- Depth Any Panoramas: A Foundation Model for Panoramic Depth EstimationXin Lin, Meixi Song, Dizhe Zhang, Wenxuan Lu et al.CVPR 2026 · 27 citations
- DA2: Depth Anything in Any DirectionHaodong Li, Wangguandong Zheng, Jing He, Yuhao Liu et al.ICLR 2026 · 23 citations
Builds on8
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Segmenter: Transformer for Semantic SegmentationRobin Strudel, Ricardo Garcia, Ivan Laptev, Cordelia SchmidICCV 2021 · 1,898 citations
- Rethinking Semantic Segmentation From a Sequence-to-Sequence Perspective With TransformersSixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu et al.CVPR 2021
Related papers
- HRDFuse: Monocular 360° Depth Estimation by Collaboratively Learning Holistic-with-Regional Depth DistributionsHao Ai, Zidong Cao, Yan-Pei Cao, Ying Shan et al.CVPR 2023
- SphereUFormer: A U-Shaped Transformer for Spherical 360 PerceptionYaniv Benny, Lior WolfCVPR 2025
- S-OmniMVS: Incorporating Sphere Geometry into Omnidirectional Stereo MatchingZisong Chen, Chunyu Lin, Lang Nie, Zhijie Shen et al.ACM MM 2023 · 8 citations
- Improving 360 Monocular Depth Estimation via Non-local Dense Prediction Transformer and Joint Supervised and Self-Supervised LearningIlwi Yun, Hyuk-Jae Lee, Chae-Eun RheeAAAI 2022 · 34 citations
- BiFuse: Monocular 360 Depth Estimation via Bi-Projection FusionFu-En Wang, Yu-Hsuan Yeh, Min Sun, Wei-Chen Chiu et al.CVPR 2020
