Zero-Shot Depth Completion with Vision-Language Model
Zhiqiang Yan, Yuan Wu, Gim Hee Lee
摘要
Vision language models (VLMs) have achieved remarkable success in semantic understanding tasks under language guidance, yet their potential for geometric perception remains largely underexplored. This paper introduces the first VLM-based depth completion framework. With almost no architectural modifications, we propose a sparse depth injection mechanism that extends the capability of VLM toward 3D perception through three key aspects: visual tokenization, textual prompt, and textual supervision. At the visual input side, sparse depth is tokenized to provide absolute scale and accurate geometric cues, alleviating the scale and camera ambiguities of RGB-only inputs. At the textual input side, a binary mask derived from sparse depth serves as a prompt, instructing the model where to complete and where to preserve. At the supervision side, the model is finetuned using text labels generated from sparse depth, requiring no ground-truth depth. Benefiting from the strong semantic priors and cross-modal expressiveness of VLM, our framework achieves superior zero-shot performance across diverse sensors, sparsity levels, and scenes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper38
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene UnderstandingMike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar 等ICCV 2021 · 被引用 633 次
相关 Paper
- DepthLM: Metric Depth from Vision Language ModelsZhipeng Cai, Ching-Feng Yeh, Hu Xu, Zhuang Liu 等ICLR 2026 · 被引用 35 次
- Vid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction-Reasoning SynergyHaijier Chen, Bo Xu, Shoujian Zhang, Haoze Liu 等ICLR 2026 · 被引用 6 次
- DenseMLLM: Standard Multimodal LLMs for Dense PredictionYi Li, Hongze Shen, Lexiang Tang, Xin Li 等ICML 2026
- SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language ModelsPingyi Chen, Yujing Lou, Shen Cao, Jinhui Guo 等NeurIPS 2025 · 被引用 24 次
- Splattalk: 3D VQA with Gaussian SplattingAnh Thai, Songyou Peng, Kyle Genova, Leonidas J. Guibas 等ICCV 2025 · 被引用 4 次
