Multi-Resolution Monocular Depth Map Fusion by Self-Supervised Gradient-Based Composition
Yaqiao Dai, Renjiao Yi, Chenyang Zhu, Hongjun He, Kai Xu
Abstract
Monocular depth estimation is a challenging problem on which deep neural networks have demonstrated great potential. However, depth maps predicted by existing deep models usually lack fine-grained details due to the convolution operations and the down-samplings in networks. We find that increasing input resolution is helpful to preserve more local details while the estimation at low resolution is more accurate globally. Therefore, we propose a novel depth map fusion module to combine the advantages of estimations with multi-resolution inputs. Instead of merging the lowand high-resolution estimations equally, we adopt the core idea of Poisson fusion, trying to implant the gradient domain of high-resolution depth into the low-resolution depth. While classic Poisson fusion requires a fusion mask as supervision, we propose a self-supervised framework based on guided image filtering. We demonstrate that this gradientbased composition performs much better at noisy immunity, compared with the state-of-the-art depth map fusion method. Our lightweight depth fusion is one-shot and runs in real-time, making our method 80X faster than a state-ofthe-art depth fusion method. Quantitative evaluations demonstrate that the proposed method can be integrated into many fully convolutional monocular depth estimation backbones with a significant performance boost, leading to state-of-theart results of detail enhancement on depth maps. Codes are released at https://github.com/yuinsky/gradient-based-depthmap-fusion .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 306c7c76-178b-446f-82ff-31bc23b2c602Cited by top-tier papers5
- Self-Distilled Depth Refinement with Noisy Poisson FusionJiaqi Li, Yiran Wang, Jinghong Zheng, Zihao Huang et al.NeurIPS 2024 · 8 citations
- Mind The Edge: Refining Depth Edges in Sparsely-Supervised Monocular Depth EstimationLior Talker, Aviad Cohen, Erez Yosef, Alexandra Dana et al.CVPR 2024 · 6 citations
- Any Resolution Any Geometry: From Multi-View To Multi-PatchWenqing Cui, Zhenyu Li, Mykola Lavreniuk, Jian Shi et al.CVPR 2026 · 2 citations
- One Look is Enough: Seamless Patchwise Refinement for Zero-Shot Monocular Depth Estimation on High-Resolution ImagesByeongjun Kwon, Munchurl KimICCV 2025 · 1 citation
- Hyden: A Hybrid Dual-Path Encoder for Monocular Geometry of High-resolution ImagesZaiwei Zhang, Marc Mapeke, Wei Ye, Rakesh Ranjan et al.ICLR 2026
Builds on8
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene UnderstandingMike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar et al.ICCV 2021 · 633 citations
- SynDeMo: Synergistic Deep Feature Alignment for Joint Learning of Depth and Ego-MotionBehzad Bozorgtabar, Mohammad Saeed Rad, Dwarikanath Mahapatra, Jean-Philippe ThiranICCV 2019 · 44 citations
- Spatial Correspondence With Generative Adversarial Network: Learning Depth From Monocular VideosZhenyao Wu, Xinyi Wu, Xiaoping Zhang, Song Wang et al.ICCV 2019 · 29 citations
Related papers
- LiteGfm: A Lightweight Self-supervised Monocular Depth Estimation Framework for Artifacts Reduction via Guided Image FilteringZhilin He, Yawei Zhang, Jingchang Mu, Xiaoyue Gu et al.ACM MM 2024 · 2 citations
- HR-Depth: High Resolution Self-Supervised Monocular Depth EstimationXiaoyang Lyu, Liang Liu, Mengmeng Wang, Xin Kong et al.AAAI 2021 · 341 citations
- R-MSFM: Recurrent Multi-Scale Feature Modulation for Monocular Depth EstimatingZhongkai Zhou, Xinnan Fan, Pengfei Shi, Yuanxue XinICCV 2021 · 150 citations
- Boosting Monocular Depth Estimation with Lightweight 3D Point FusionLam Huynh, Phong Nguyen, Jirí Matas, Esa Rahtu et al.ICCV 2021 · 32 citations
- Unsupervised High-Resolution Depth Learning From Videos With Dual NetworksJunsheng Zhou, Yuwang Wang, Kaihuai Qin, Wenjun ZengICCV 2019 · 77 citations
