WorDepth: Variational Language Prior for Monocular Depth Estimation
Ziyao Zeng, Daniel Wang, Fengyu Yang, Hyoungseob Park, Stefano Soatto, Dong Lao, Alex Wong
摘要
Three-dimensional (3D) reconstruction from a single image is an ill-posed problem with inherent ambiguities, i. e. scale. Predicting a 3D scene from text descriptionis) is similarly ill-posed, i. e. spatial arrangements of objects described. We investigate the question of whether two inher-ently ambiguous modalities can be used in conjunction to produce metric-scaled reconstructions. To test this, we fo-cus on monocular depth estimation, the problem of predicting a dense depth map from a single image, but with an additional text caption describing the scene. To this end, we begin by encoding the text caption as a mean and standard deviation; using a variational framework, we learn the distribution of the plausible metric reconstructions of 3D scenes corresponding to the text captions as a prior. To “select” a specific reconstruction or depth map, we encode the given image through a conditional sampler that samples from the latent space of the variational text encoder, which is then decoded to the output depth map. Our approach is trained alternatingly between the text and image branches: in one optimization step, we predict the mean and standard deviation from the text description and sample from a standard Gaussian, and in the other, we sample using a (image) conditional sampler. Once trained, we directly predict depth from the encoded text using the conditional sampler. We demonstrate our approach on indoor (NYUv2) and out-door (KITTI) scenarios, where we show that language can consistently improve performance in both. Code: https://github.com/Adonis-galaxy/WorDepth.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Dual Prototype Evolving for Test-Time Generalization of Vision-Language ModelsCe Zhang, Simon Stepputtis, Katia P. Sycara, Yaqi XieNeurIPS 2024 · 被引用 57 次
- Jasmine: Harnessing Diffusion Prior for Self-supervised Depth EstimationJiyuan Wang, Chunyu Lin, Cheng Guan, Lang Nie 等NeurIPS 2025 · 被引用 26 次
- RSA: Resolving Scale Ambiguities in Monocular Depth Estimators through Language DescriptionsZiyao Zeng, Yangchao Wu, Hyoungseob Park, Daniel Wang 等NeurIPS 2024 · 被引用 26 次
- Metric from Human: Zero-shot Monocular Metric Depth Estimation via Test-time AdaptationYizhou Zhao, Hengwei Bian, Kaihua Chen, Pengliang Ji 等NeurIPS 2024 · 被引用 16 次
- Any3D-VLA: Enhancing VLA Robustness via Diverse Point CloudsXianzhe Fan, Shengliang Deng, Xiaoyang Wu, Yuxiang Lu 等ICML 2026 · 被引用 9 次
它引用的顶会 Paper41
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- Vision-Language Embodiment for Monocular Depth EstimationJinchang Zhang, Guoyu LuCVPR 2025
- Iris: Integrating Language into Diffusion-based Monocular Depth EstimationZiyao Zeng, Jingcheng Ni, Daniel Wang, Patrick Rim 等CVPR 2026
- Towards Zero-Shot Scale-Aware Monocular Depth EstimationVitor Guizilini, Igor Vasiljevic, Dian Chen, Rares Ambrus 等ICCV 2023 · 被引用 129 次
- TPDepth: Leveraging Text Prompts with ControlNet to Boost Diffusion-based Depth EstimationYu Liu, Kun Sun, Chang Tang, Yuhua Qian 等ACM MM 2025 · 被引用 2 次
- Text2Room: Extracting Textured 3D Meshes from 2D Text-to-Image ModelsLukas Höllein, Ang Cao, Andrew Owens, Justin Johnson 等ICCV 2023 · 被引用 292 次
