TPDepth: Leveraging Text Prompts with ControlNet to Boost Diffusion-based Depth Estimation
Yu Liu, Kun Sun, Chang Tang, Yuhua Qian, Xin Li
摘要
Recent diffusion-based methods have shown strong ability in the depth estimation task, but they largely overlook the rich textual priors embedded in pretrained diffusion models that can enhance both performance and robustness in diverse scenes. In this paper, we propose TPDepth, a diffusion-based, affine-invariant monocular depth estimator that incorporates textual semantics via a Text-Prompted ControlNet. While directly injecting text into the diffusion U-Net can cause the network to over-attend to local semantic cues and compromise global structural modeling, TPDepth processes textual features through a separate ControlNet branch, allowing semantic information to be incorporated without disrupting the spatial reasoning pipeline. Prompt-conditioned features are modulated by an Adaptive Control Scale Module(ACSM) and injected into decoder of the diffusion UNet with skip connections. The model is fine-tuned with a fixed timestep for deterministic prediction. TPDepth achieves state-of-the-art results on NYUv2, KITTI, and ScanNet, and demonstrates competitive performance on two additional zero-shot benchmarks using only 61K training images. Code and models can be found on our https://github.com/Lioely/TPDepth project page.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- SGAT: Learning Feature Matching with Singularity-enhanced Graph Attention NetworkYizhuo Zhang, Kun Sun, Chang Tang, Yuanyuan Liu 等AAAI 2026
- GeoCoBox: Box-supervised 3D Tumor Segmentation via Geometric Co-embeddingTianzhong Lan, Zhang Yi, Xiuyuan Xu, Min ZhuAAAI 2026
相关 Paper
- OGDepth: Leveraging Object Guidance in Diffusion Models for Enhanced Monocular Depth EstimationWenzheng Yang, Songwei Pei, Bingfeng Liu, Qian Li 等ACM MM 2025
- Unleashing Text-to-Image Diffusion Models for Visual PerceptionWenliang Zhao, Yongming Rao, Zuyan Liu, Benlin Liu 等ICCV 2023 · 被引用 327 次
- Text-Image Alignment for Diffusion-Based PerceptionNeehar Kondapaneni, Markus Marks, Manuel Knott, Rogério Guimarães 等CVPR 2024
- BetterDepth: Plug-and-Play Diffusion Refiner for Zero-Shot Monocular Depth EstimationXiang Zhang, Bingxin Ke, Hayko Riemenschneider, Nando Metzger 等NeurIPS 2024 · 被引用 32 次
- Iris: Integrating Language into Diffusion-based Monocular Depth EstimationZiyao Zeng, Jingcheng Ni, Daniel Wang, Patrick Rim 等CVPR 2026
