Vision-Language Embodiment for Monocular Depth Estimation
Jinchang Zhang, Guoyu Lu
Abstract
Depth estimation is a core problem in robotic perception and vision tasks, but 3D reconstruction from a single image presents inherent uncertainties. Current depth estimation models primarily rely on inter-image relationships for supervised training, often overlooking the intrinsic information provided by the camera itself. We propose a method that embodies the camera model and its physical characteristics into a deep learning model, computing embodied scene depth through real-time interactions with road environments. The model can calculate embodied scene depth in real-time based on immediate environmental changes using only the intrinsic properties of the camera, without any additional equipment. By combining embodied scene depth with RGB image features, the model gains a comprehensive perspective on both geometric and visual details. Additionally, we incorporate text descriptions containing environmental content and depth information as priors for scene understanding, enriching the model's perception of objects. This integration of image and language -two inherently ambiguous modalities -leverages their complementary strengths for monocular depth estimation. The real-time nature of the embodied language and depth prior model ensures that the model can continuously adjust its perception and behavior in dynamic environments. Experimental results show that the embodied depth estimation method enhances model performance across different scenes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e3e21fef-ff13-4b65-8ab0-dd31bad09176Cited by top-tier papers3
- Any3D-VLA: Enhancing VLA Robustness via Diverse Point CloudsXianzhe Fan, Shengliang Deng, Xiaoyang Wu, Yuxiang Lu et al.ICML 2026 · 9 citations
- AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation ModelsZheda Mai, Arpita Chowdhury, Zihe Wang, Sooyoung Jeon et al.CVPR 2026 · 7 citations
- BlackMirror: Black-Box Backdoor Detection for Text-to-Image Models via Instruction-Response DeviationFeiran Li, Qianqian Xu, Shilong Bao, Zhiyong Yang et al.CVPR 2026 · 2 citations
Builds on19
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat et al.CVPR 2022 · 3,348 citations
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
- Enforcing Geometric Constraints of Virtual Normal for Depth PredictionWei Yin, Yifan Liu, Chunhua Shen, Youliang YanICCV 2019 · 487 citations
- P3Depth: Monocular Depth Estimation with a Piecewise Planarity PriorVaishakh Patil, Christos Sakaridis, Alexander Liniger, Luc Van GoolCVPR 2022 · 144 citations
Related papers
- WorDepth: Variational Language Prior for Monocular Depth EstimationZiyao Zeng, Daniel Wang, Fengyu Yang, Hyoungseob Park et al.CVPR 2024 · 20 citations
- MonoDTR: Monocular 3D Object Detection with Depth-Aware TransformerKuan-Chih Huang, Tsung-Han Wu, Hung-Ting Su, Winston H. HsuCVPR 2022 · 199 citations
- GEDepth: Ground Embedding for Monocular Depth EstimationXiaodong Yang, Zhuang Ma, Zhiyu Ji, Zhe RenICCV 2023 · 40 citations
- PromptDepth: Efficient and Promptable Geometric 3D Vision Model for Embodied IntelligenceXianyun Wang, Jiaxu Miao, Tian Xu, Siyuan Wang et al.CVPR 2026
- Holistic 3D Human and Scene Mesh Estimation From Single View ImagesZhenzhen Weng, Serena YeungCVPR 2021
