ECoDepth: Effective Conditioning of Diffusion Models for Monocular Depth Estimation
Suraj Patni, Aradhye Agarwal, Chetan Arora
摘要
In the absence of parallax cues, a learning based single image depth estimation (SIDE) model relies heavily on shading and contextual cues in the image. While this simplicity is attractive, it is necessary to train such models on large and varied datasets, which are difficult to capture. It has been shown that using embeddings from pretrained foundational models, such as CLIP, improves zero shot transfer in several applications. Taking inspiration from this, in our paper we explore the use of global image priors generated from a pretrained ViT model to provide more detailed contextual information. We argue that the embedding vector from a ViT model, pretrained on a large dataset, captures greater relevant information for SIDE than the usual route of generating pseudo image captions, followed by CLIP based text embeddings. Based on this idea, we propose a new SIDE model using a diffusion backbone which is conditioned on ViT embeddings. Our proposed design establishes a new state-of-the-art (SOTA) for SIDE on NYU Depth v2 dataset, achieving Abs Rel error of 0.059(14% improvement) compared to 0.069 by the current SOTA (VPD). And on KITTI dataset, achieving Sq Rel error of 0.139 (2% improvement) compared to 0.142 by the current SOTA (GED). For zero shot transfer with a model trained on NYU Depth v2, we report mean relative improvement of (20%, 23%,81%, 25%) over NeWCRF on (Sun-RGBD, iBimsl, DIODE, HyperSim) datasets, compared to (16%, 18%, 45%, 9%) by ZoEDepth. The code is available in our project page.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- A General Protocol to Probe Large Vision Models for 3D Physical UnderstandingGuanqi Zhan, Chuanxia Zheng, Weidi Xie, Andrew ZissermanNeurIPS 2024 · 被引用 37 次
- Digging into Contrastive Learning for Robust Depth Estimation with Diffusion ModelsJiyuan Wang, Chunyu Lin, Lang Nie, Kang Liao 等ACM MM 2024 · 被引用 7 次
- un2CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIPYinqi Li, Jiahe Zhao, Hong Chang, Ruibing Hou 等NeurIPS 2025 · 被引用 6 次
- Test-Time Prompt Tuning for Zero-Shot Depth CompletionChanhwi Jeong, Inhwan Bae, Jin-Hwi Park, Hae-Gon JeonICCV 2025 · 被引用 3 次
- A Simple Yet Mighty Hartley Diffusion Versatilist for Generalizable Dense Vision TasksQi Bi, Jingjun Yi, Huimin Huang, Hao Zheng 等ICCV 2025 · 被引用 3 次
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
相关 Paper
- Towards Zero-Shot Scale-Aware Monocular Depth EstimationVitor Guizilini, Igor Vasiljevic, Dian Chen, Rares Ambrus 等ICCV 2023 · 被引用 129 次
- Hybrid-Grained Feature Aggregation with Coarse-to-Fine Language Guidance for Self-Supervised Monocular Depth EstimationWenyao Zhang, Hongsi Liu, Bohan Li, Jiawei He 等ICCV 2025 · 被引用 2 次
- EZSR: Event-based Zero-Shot RecognitionYan Yang, Liyuan Pan, Dongxu Li, Liu LiuCVPR 2025
- TPDepth: Leveraging Text Prompts with ControlNet to Boost Diffusion-based Depth EstimationYu Liu, Kun Sun, Chang Tang, Yuhua Qian 等ACM MM 2025 · 被引用 2 次
- Mask3D: Pretraining 2D Vision Transformers by Learning Masked 3D PriorsJi Hou, Xiaoliang Dai, Zijian He, Angela Dai 等CVPR 2023
