OGDepth: Leveraging Object Guidance in Diffusion Models for Enhanced Monocular Depth Estimation
Wenzheng Yang, Songwei Pei, Bingfeng Liu, Qian Li, Shangguang Wang
摘要
Monocular depth estimation stands as a fundamental pursuit in computer vision. Recently, some methods have attempted to introduce the text-to-image diffusion model into the domain of monocular depth estimation and achieved impressive results. However, these methods typically employ pre-defined templates as text prompts to guide the learning of denoising networks, resulting in limited flexibility and scalability. In this paper, we propose OGDepth, a diffusion-based monocular depth estimation network with object prompts generated by taking advantage of the object detection information from the scene. Specifically, we design an Object Prompt Module (OPM) to encode the object detection information into prompts that are more closely aligned with the image content, offering richer contextual information while circumventing the monotony and redundancy inherent in template-generated prompts. Moreover, we employ bounding box information for each object to filter and localize objects, enabling the model to grasp relative positional information within the scene. This facilitates the creation of a more precise depth map. Additionally, we design a Global-Local Interaction Decoder (GLID) to facilitate the mutual exchange of features at different scales, enabling efficient feature fusion. Our approach underwent rigorous experiments across multiple datasets, with results showcasing its state-of-the-art performance. Notably, on the KITTI dataset, our model achieves an RMSE of 1.967 and a REL of 0.047, and both metrics are the best among all compared methods. On the NYU Depth V2 dataset, our method achieves an RMSE score of 0.221, representing a notable 12.9% enhancement compared to the baseline method (VPD).
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- TPDepth: Leveraging Text Prompts with ControlNet to Boost Diffusion-based Depth EstimationYu Liu, Kun Sun, Chang Tang, Yuhua Qian 等ACM MM 2025 · 被引用 2 次
- MonoDiff: Monocular 3D Object Detection and Pose Estimation with Diffusion ModelsYasiru Ranasinghe, Deepti Hegde, Vishal M. PatelCVPR 2024 · 被引用 21 次
- MonoDTR: Monocular 3D Object Detection with Depth-Aware TransformerKuan-Chih Huang, Tsung-Han Wu, Hung-Ting Su, Winston H. HsuCVPR 2022 · 被引用 199 次
- MoGDE: Boosting Mobile Monocular 3D Object Detection with Ground Depth EstimationYunsong Zhou, Quan Liu, Hongzi Zhu, Yunzhe Li 等NeurIPS 2022 · 被引用 23 次
- MonoDETR: Depth-guided Transformer for Monocular 3D Object DetectionRenrui Zhang, Han Qiu, Tai Wang, Ziyu Guo 等ICCV 2023 · 被引用 175 次
