WorDepth: Variational Language Prior for Monocular Depth Estimation
Ziyao Zeng, Daniel Wang, Fengyu Yang, Hyoungseob Park, Stefano Soatto, Dong Lao, Alex Wong
Abstract
Three-dimensional (3D) reconstruction from a single image is an ill-posed problem with inherent ambiguities, i. e. scale. Predicting a 3D scene from text descriptionis) is similarly ill-posed, i. e. spatial arrangements of objects described. We investigate the question of whether two inher-ently ambiguous modalities can be used in conjunction to produce metric-scaled reconstructions. To test this, we fo-cus on monocular depth estimation, the problem of predicting a dense depth map from a single image, but with an additional text caption describing the scene. To this end, we begin by encoding the text caption as a mean and standard deviation; using a variational framework, we learn the distribution of the plausible metric reconstructions of 3D scenes corresponding to the text captions as a prior. To “select” a specific reconstruction or depth map, we encode the given image through a conditional sampler that samples from the latent space of the variational text encoder, which is then decoded to the output depth map. Our approach is trained alternatingly between the text and image branches: in one optimization step, we predict the mean and standard deviation from the text description and sample from a standard Gaussian, and in the other, we sample using a (image) conditional sampler. Once trained, we directly predict depth from the encoded text using the conditional sampler. We demonstrate our approach on indoor (NYUv2) and out-door (KITTI) scenarios, where we show that language can consistently improve performance in both. Code: https://github.com/Adonis-galaxy/WorDepth.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 580857e5-599f-43f4-85d9-1d423a2c0ee8Cited by top-tier papers22
- Dual Prototype Evolving for Test-Time Generalization of Vision-Language ModelsCe Zhang, Simon Stepputtis, Katia P. Sycara, Yaqi XieNeurIPS 2024 · 57 citations
- Jasmine: Harnessing Diffusion Prior for Self-supervised Depth EstimationJiyuan Wang, Chunyu Lin, Cheng Guan, Lang Nie et al.NeurIPS 2025 · 26 citations
- RSA: Resolving Scale Ambiguities in Monocular Depth Estimators through Language DescriptionsZiyao Zeng, Yangchao Wu, Hyoungseob Park, Daniel Wang et al.NeurIPS 2024 · 26 citations
- Metric from Human: Zero-shot Monocular Metric Depth Estimation via Test-time AdaptationYizhou Zhao, Hengwei Bian, Kaihua Chen, Pengliang Ji et al.NeurIPS 2024 · 16 citations
- Any3D-VLA: Enhancing VLA Robustness via Diverse Point CloudsXianzhe Fan, Shengliang Deng, Xiaoyang Wu, Yuxiang Lu et al.ICML 2026 · 9 citations
Builds on41
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Vision-Language Embodiment for Monocular Depth EstimationJinchang Zhang, Guoyu LuCVPR 2025
- Iris: Integrating Language into Diffusion-based Monocular Depth EstimationZiyao Zeng, Jingcheng Ni, Daniel Wang, Patrick Rim et al.CVPR 2026
- Towards Zero-Shot Scale-Aware Monocular Depth EstimationVitor Guizilini, Igor Vasiljevic, Dian Chen, Rares Ambrus et al.ICCV 2023 · 129 citations
- TPDepth: Leveraging Text Prompts with ControlNet to Boost Diffusion-based Depth EstimationYu Liu, Kun Sun, Chang Tang, Yuhua Qian et al.ACM MM 2025 · 2 citations
- Text2Room: Extracting Textured 3D Meshes from 2D Text-to-Image ModelsLukas Höllein, Ang Cao, Andrew Owens, Justin Johnson et al.ICCV 2023 · 292 citations
