PatchRefiner V2: Fast and Lightweight Real-Domain High-Resolution Metric Depth Estimation
Zhenyu Li, Wenqing Cui, Shariq Farooq Bhat, Peter Wonka
摘要
While current high-resolution depth estimation methods achieve strong results, they often suffer from computational inefficiencies due to reliance on heavyweight models and multiple inference steps, increasing inference time. To address this, we introduce PatchRefiner V2 (PRV2), which replaces heavy refiner models with lightweight encoders. This reduces model size and inference time but introduces noisy features. To overcome this, we propose a Coarse-to-Fine (C2F) module with a Guided Denoising Unit for refining and denoising the refiner features and a Noisy Pretraining strategy to pretrain the refiner branch to fully exploit the potential of the lightweight refiner branch. Additionally, we introduce a Scale-and-Shift Invariant Gradient Matching (SSIGM) loss to enhance synthetic-to-real domain transfer. PRV2 outperforms state-of-the-art depth estimation methods on UnrealStereo4K in both accuracy and speed, using fewer parameters and faster inference. It also shows improved depth boundary delineation on real-world datasets like CityScape, ScanNet++, and KITTI, demonstrating its versatility across domains.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Any Resolution Any Geometry: From Multi-View To Multi-PatchWenqing Cui, Zhenyu Li, Mykola Lavreniuk, Jian Shi 等CVPR 2026 · 被引用 2 次
- LiteSense: Lifting Lightweight ToF with RGB for High-Resolution Metric Depth EstimationYusheng Li, Lizhi LOU, Yan Tang, Zekai Miao 等CVPR 2026
它引用的顶会 Paper31
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- Self-Distilled Depth Refinement with Noisy Poisson FusionJiaqi Li, Yiran Wang, Jinghong Zheng, Zihao Huang 等NeurIPS 2024 · 被引用 8 次
- One Look is Enough: Seamless Patchwise Refinement for Zero-Shot Monocular Depth Estimation on High-Resolution ImagesByeongjun Kwon, Munchurl KimICCV 2025 · 被引用 1 次
- A Decomposition Model for Stereo MatchingChengtang Yao, Yunde Jia, Huijun Di, Pengxiang Li 等CVPR 2021
- PatchFusion: An End-to-End Tile-Based Framework for High-Resolution Monocular Metric Depth EstimationZhenyu Li, Shariq Farooq Bhat, Peter WonkaCVPR 2024
- GraftNet: Towards Domain Generalized Stereo Matching with a Broad-Spectrum and Task-Oriented FeatureBiyang Liu, Huimin Yu, Guodong QiCVPR 2022 · 被引用 52 次
