PatchFusion: An End-to-End Tile-Based Framework for High-Resolution Monocular Metric Depth Estimation
Zhenyu Li, Shariq Farooq Bhat, Peter Wonka
Abstract
Abstract Single image depth estimation is a foundational task in computer vision and generative modeling. However, prevailing depth estimation models grapple with accommodating the increasing resolutions commonplace in today's consumer cameras and devices. Existing highresolution strategies show promise, but they often face limitations, ranging from error propagation to the loss of highfrequency details. We present PatchFusion, a novel tilebased framework with three key components to improve the current state of the art: (1) A patch-wise fusion network that fuses a globally-consistent coarse prediction with finer, inconsistent tiled predictions via high-level feature guidance, (2) A Global-to-Local (G2L) module that adds vital context to the fusion network, discarding the need for patch selection heuristics, and (3) A Consistency-Aware Training (CAT) and Inference (CAI) approach, emphasizing patch overlap consistency and thereby eradicating the necessity for post-processing. Experiments on UnrealStereo4K, MVS-Synth, and Middleburry 2014 demonstrate that our framework can generate high-resolution depth maps with intricate details. PatchFusion is independent of the base model for depth estimation. Notably, our framework built on top of SOTA ZoeDepth brings improvements for a total of 17.3% and 29.4% in terms of the root mean squared error (RMSE) on UnrealStereo4K and MVS-Synth, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 417d2dc4-4958-43b4-911a-d1a35c11bd52Cited by top-tier papers23
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp DetailsRuicheng Wang, Sicheng Xu, Yue Dong, Yu Deng et al.NeurIPS 2025 · 308 citations
- Pixel-Perfect Depth with Semantics-Prompted Diffusion TransformersGangwei Xu, Haotong Lin, Hongcheng Luo, Xianqi Wang et al.NeurIPS 2025 · 58 citations
- InfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit FieldsHao Yu, Haotong Lin, Jiawei Wang, Jiaxin Li et al.CVPR 2026 · 19 citations
- Depth Pro: Sharp Monocular Metric Depth in Less Than a SecondAlexey Bochkovskiy, Amaël Delaunoy, Hugo Germain, Marcel Santos et al.ICLR 2025 · 15 citations
Builds on19
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Unleashing Text-to-Image Diffusion Models for Visual PerceptionWenliang Zhao, Yongming Rao, Zuyan Liu, Benlin Liu et al.ICCV 2023 · 327 citations
Related papers
- One Look is Enough: Seamless Patchwise Refinement for Zero-Shot Monocular Depth Estimation on High-Resolution ImagesByeongjun Kwon, Munchurl KimICCV 2025 · 1 citation
- Self-Distilled Depth Refinement with Noisy Poisson FusionJiaqi Li, Yiran Wang, Jinghong Zheng, Zihao Huang et al.NeurIPS 2024 · 8 citations
- Hyper-Depth: Hypergraph-Based Multi-Scale Representation Fusion for Monocular Depth EstimationLin Bie, Siqi Li, Yifan Feng, Yue GaoICCV 2025
- Any Resolution Any Geometry: From Multi-View To Multi-PatchWenqing Cui, Zhenyu Li, Mykola Lavreniuk, Jian Shi et al.CVPR 2026 · 2 citations
- Multi-Resolution Monocular Depth Map Fusion by Self-Supervised Gradient-Based CompositionYaqiao Dai, Renjiao Yi, Chenyang Zhu, Hongjun He et al.AAAI 2023 · 8 citations
