TPDepth: Leveraging Text Prompts with ControlNet to Boost Diffusion-based Depth Estimation
Yu Liu, Kun Sun, Chang Tang, Yuhua Qian, Xin Li
Abstract
Recent diffusion-based methods have shown strong ability in the depth estimation task, but they largely overlook the rich textual priors embedded in pretrained diffusion models that can enhance both performance and robustness in diverse scenes. In this paper, we propose TPDepth, a diffusion-based, affine-invariant monocular depth estimator that incorporates textual semantics via a Text-Prompted ControlNet. While directly injecting text into the diffusion U-Net can cause the network to over-attend to local semantic cues and compromise global structural modeling, TPDepth processes textual features through a separate ControlNet branch, allowing semantic information to be incorporated without disrupting the spatial reasoning pipeline. Prompt-conditioned features are modulated by an Adaptive Control Scale Module(ACSM) and injected into decoder of the diffusion UNet with skip connections. The model is fine-tuned with a fixed timestep for deterministic prediction. TPDepth achieves state-of-the-art results on NYUv2, KITTI, and ScanNet, and demonstrates competitive performance on two additional zero-shot benchmarks using only 61K training images. Code and models can be found on our https://github.com/Lioely/TPDepth project page.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers2
- SGAT: Learning Feature Matching with Singularity-enhanced Graph Attention NetworkYizhuo Zhang, Kun Sun, Chang Tang, Yuanyuan Liu et al.AAAI 2026
- GeoCoBox: Box-supervised 3D Tumor Segmentation via Geometric Co-embeddingTianzhong Lan, Zhang Yi, Xiuyuan Xu, Min ZhuAAAI 2026
Related papers
- OGDepth: Leveraging Object Guidance in Diffusion Models for Enhanced Monocular Depth EstimationWenzheng Yang, Songwei Pei, Bingfeng Liu, Qian Li et al.ACM MM 2025
- Unleashing Text-to-Image Diffusion Models for Visual PerceptionWenliang Zhao, Yongming Rao, Zuyan Liu, Benlin Liu et al.ICCV 2023 · 327 citations
- Text-Image Alignment for Diffusion-Based PerceptionNeehar Kondapaneni, Markus Marks, Manuel Knott, Rogério Guimarães et al.CVPR 2024
- BetterDepth: Plug-and-Play Diffusion Refiner for Zero-Shot Monocular Depth EstimationXiang Zhang, Bingxin Ke, Hayko Riemenschneider, Nando Metzger et al.NeurIPS 2024 · 32 citations
- Iris: Integrating Language into Diffusion-based Monocular Depth EstimationZiyao Zeng, Jingcheng Ni, Daniel Wang, Patrick Rim et al.CVPR 2026
