HRDFuse: Monocular 360° Depth Estimation by Collaboratively Learning Holistic-with-Regional Depth Distributions
Hao Ai, Zidong Cao, Yan-Pei Cao, Ying Shan, Lin Wang
Abstract
Depth estimation from a monocular 360 • image is a burgeoning problem owing to its holistic sensing of a scene. Recently, some methods, e.g., OmniFusion, have applied the tangent projection (TP) to represent a 360 • image and predicted depth values via patch-wise regressions, which are merged to get a depth map with equirectangular projection (ERP) format. However, these methods suffer from 1) non-trivial process of merging plenty of patches; 2) capturing less holistic-with-regional contextual information by directly regressing the depth value of each pixel. In this paper, we propose a novel framework, HRDFuse, that subtly combines the potential of convolutional neural networks (CNNs) and transformers by collaboratively learning the holistic contextual information from the ERP and the regional structural information from the TP. Firstly, we propose a spatial feature alignment (SFA) module that learns feature similarities between the TP and ERP to aggregate the TP features into a complete ERP feature map in a pixelwise manner. Secondly, we propose a collaborative depth distribution classification (CDDC) module that learns the holistic-with-regional histograms capturing the ERP and TP depth distributions. As such, the final depth values can be predicted as a linear combination of histogram bin centers. Lastly, we adaptively combine the depth predictions from ERP and TP to obtain the final depth map. Extensive experiments show that our method predicts more smooth and accurate depth results while achieving favorably better results than the SOTA methods.
For videos, code, demo and more information, you can visit https://VLIS2022.github.io/HRDFuse/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Dream360: Diverse and Immersive Outdoor Virtual Scene Creation via Transformer-Based 360° Image OutpaintingHao Ai, Zidong Cao, Haonan Lu, Chen Chen et al.IEEE VR 2024 · 12 citations
- UniDAC: Universal Metric Depth Estimation for Any CameraGirish Chandar Ganesan, Yuliang Guo, Liu Ren, Xiaoming LiuCVPR 2026 · 8 citations
- Depth Any Camera: Zero-Shot Metric Depth Estimation from Any CameraYuliang Guo, Sparsh Garg, S. Mahdi H. Miangoleh, Xinyu Huang et al.CVPR 2025
- SO(3)-Equivariant ViT-Adapter for Data-Efficient Zero-Shot Sim-to-Real Indoor Panoramic Depth EstimationZiyan He, Qiudan Zhang, Lin Ma, Xu WangCVPR 2026
- Elite360D: Towards Efficient 360 Depth Estimation via Semantic- and Distance-Aware Bi-Projection FusionHao Ai, Lin WangCVPR 2024
Builds on16
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan et al.CVPR 2022 · 1,603 citations
- CMT: Convolutional Neural Networks Meet Vision TransformersJianyuan Guo, Kai Han, Han Wu, Yehui Tang et al.CVPR 2022 · 839 citations
- Conformer: Local Features Coupling Global Representations for Visual RecognitionZhiliang Peng, Wei Huang, Shanzhi Gu, Lingxi Xie et al.ICCV 2021 · 723 citations
- Mobile-Former: Bridging MobileNet and TransformerYinpeng Chen, Xiyang Dai, Dongdong Chen, Mengchen Liu et al.CVPR 2022 · 600 citations
Related papers
- OmniFusion: 360 Monocular Depth Estimation via Geometry-Aware FusionYuyan Li, Yuliang Guo, Zhixin Yan, Xinyu Huang et al.CVPR 2022 · 79 citations
- 360MonoDepth: High-Resolution 360° Monocular Depth EstimationManuel Rey-Area, Mingze Yuan, Christian RichardtCVPR 2022 · 80 citations
- BiFuse: Monocular 360 Depth Estimation via Bi-Projection FusionFu-En Wang, Yu-Hsuan Yeh, Min Sun, Wei-Chen Chiu et al.CVPR 2020
- 360 Depth Estimation in the Wild - the Depth360 Dataset and the SegFuse NetworkQi Feng, Hubert P. H. Shum, Shigeo MorishimaIEEE VR 2022 · 24 citations
- Improving 360 Monocular Depth Estimation via Non-local Dense Prediction Transformer and Joint Supervised and Self-Supervised LearningIlwi Yun, Hyuk-Jae Lee, Chae-Eun RheeAAAI 2022 · 34 citations
