DepthLM: Metric Depth from Vision Language Models
Zhipeng Cai, Ching-Feng Yeh, Hu Xu, Zhuang Liu, Gregory P. Meyer, Xinjie Lei, Changsheng Zhao, Shang-Wen Li, Vikas Chandra, Yangyang Shi
Abstract
Vision language models (VLMs) can flexibly address various vision tasks through text interactions. Although successful in semantic understanding, state-of-the-art VLMs including GPT-5 still struggle in understanding 3D from 2D inputs. On the other hand, expert pure vision models achieve super-human accuracy in metric depth estimation, a key 3D understanding task. However, they require task-specific architectures and losses. Such difference motivates us to ask: Can VLMs reach expert-level accuracy without architecture or loss change? We take per-pixel metric depth estimation as the representative task and show that the answer is yes! Surprisingly, comprehensive analysis shows that text-based supervised-finetuning with sparse labels is sufficient for VLMs to unlock strong 3D understanding, no dense prediction head or complex regression/regularization loss is needed. The bottleneck for VLMs lies actually in pixel reference and cross-dataset camera ambiguity, which we address through visual prompting and intrinsic-conditioned augmentation. With much smaller models, our method DepthLM surpasses the accuracy of most advanced VLMs by over 2x, making VLMs for the first time comparable with pure vision models. Interestingly, without explicit enforcement during training, VLMs trained with DepthLM naturally avoids over-smoothing, having much fewer flying points at boundary regions than pure vision models. The simplicity of DepthLM also enables a single VLM to cover various 3D tasks beyond metric depth. Our code and model will be released at the link below.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RLSiyi Chen, Mikaela Angelina Uy, Chan Hee Song, Faisal Ladhak et al.CVPR 2026 · 24 citations
- DepthFocus: Controllable Depth Estimation for See-Through Scenesjunhong min, Jimin Kim, Minwook Kim, Cheol-Hui Min et al.CVPR 2026 · 4 citations
- WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian NavigationRafi Ibn Sultan, Hui Zhu, Xiangyu Zhou, Chengyin Li et al.CVPR 2026 · 4 citations
- Linking Perception, Confidence and Accuracy in MLLMsYuetian Du, Yucheng Wang, Rongyu Zhang, Zhijie Xu et al.CVPR 2026 · 2 citations
- Zero-Shot Depth Completion with Vision-Language ModelZhiqiang Yan, Yuan Wu, Gim Hee LeeCVPR 2026 · 1 citation
Builds on13
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- ScanNet++: A High-Fidelity Dataset of 3D Indoor ScenesChandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, Angela DaiICCV 2023 · 659 citations
- SpatialRGPT: Grounded Spatial Reasoning in Vision-Language ModelsAn-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo et al.NeurIPS 2024 · 412 citations
- Metric3D: Towards Zero-shot Metric 3D Prediction from A Single ImageWei Yin, Chi Zhang, Hao Chen, Zhipeng Cai et al.ICCV 2023 · 388 citations
Related papers
- DenseMLLM: Standard Multimodal LLMs for Dense PredictionYi Li, Hongze Shen, Lexiang Tang, Xin Li et al.ICML 2026
- Vid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction-Reasoning SynergyHaijier Chen, Bo Xu, Shoujian Zhang, Haoze Liu et al.ICLR 2026 · 6 citations
- MonoVLM: Monocular 3D Visual Grounding with Vision Language ModelsHuaizhi Qu, Hossein Nourkhiz Mahjoub, Vaishnav Tadiparthi, Kwonjoon Lee et al.CVPR 2026
- InteractVLM: 3D Interaction Reasoning from 2D Foundational ModelsSai Kumar Dwivedi, Dimitrije Antic, Shashank Tripathi, Omid Taheri et al.CVPR 2025
- Leveraging VLM-Based Pipelines to Annotate 3D ObjectsRishabh Kabra, Loic Matthey, Alexander Lerchner, Niloy J. MitraICML 2024 · 10 citations
