SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models
Pingyi Chen, Yujing Lou, Shen Cao, Jinhui Guo, Lubin Fan, Yue Wu, Lin F. Yang, Lizhuang Ma, Jieping Ye
Abstract
While vision language models (VLMs) excel in 2D semantic visual understanding, their ability to quantitatively reason about 3D spatial relationships remains underexplored due to the deficiency of spatial representation ability of 2D images. In this paper, we analyze the problem hindering VLMs' spatial understanding abilities and propose SD-VLM, a novel framework that significantly enhances fundamental spatial perception abilities of VLMs through two key contributions:
(1) propose Massive Spatial Measuring and Understanding (MSMU) dataset with precise spatial annotations, and (2) introduce a simple depth positional encoding method strengthening VLMs' spatial awareness. MSMU dataset includes massive quantitative spatial tasks with 700K QA pairs, 2.5M physical numerical annotations, and 10K chain-of-thought augmented samples. We have trained SD-VLM, a strong generalist VLM which shows superior quantitative spatial measuring and understanding capability. SD-VLM not only achieves state-of-the-art performance on our proposed MSMU-Bench, but also shows spatial generalization abilities on other spatial understanding benchmarks including Q-Spatial and SpatialRGPT-Bench. Extensive experiments demonstrate that SD-VLM outperforms GPT-4o and Intern-VL3-78B by 26.91% and 25.56% respectively on MSMU-Bench. Code and models are released at https://github.com/cpystan/SD-VLM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c58e775d-a5ea-43c5-89b0-6f77cfd03174Cited by top-tier papers5
- 4D-RGPT: Toward Region-level 4D Understanding via Perceptual DistillationChiao-An Yang, Ryo Hachiuma, Sifei Liu, Subhashree Radhakrishnan et al.CVPR 2026 · 2 citations
- HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language ModelsHuizhi Liang, Yichao Shen, Yu Deng, Sicheng Xu et al.CVPR 2026 · 2 citations
- From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language ModelsMasanari Oi, Koki Maeda, Ryuto Koike, Daisuke Oba et al.ICML 2026 · 2 citations
- OmniLottie: Generating Vector Animations via Parameterized Lottie TokensYiying Yang, Wei Cheng, Sijin Chen, Honghao Fu et al.CVPR 2026 · 2 citations
- Keep it SymPL: Symbolic Projective Layout for Allocentric Spatial Reasoning in Vision-Language ModelsJaeyun Jang, Seunghui Shin, Taeho Park, Hyoseok HwangCVPR 2026 · 1 citation
Builds on30
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
- 3D-LLM: Injecting the 3D World into Large Language ModelsYining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng et al.NeurIPS 2023 · 662 citations
Related papers
- SpatialRGPT: Grounded Spatial Reasoning in Vision-Language ModelsAn-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo et al.NeurIPS 2024 · 412 citations
- Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language ModelsRunsen Xu, Weiyao Wang, Hao Tang, Xingyu Chen et al.CVPR 2026 · 64 citations
- MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMsErik A. Daxberger, Nina Wenzel, David Griffiths, Haiming Gang et al.ICCV 2025 · 10 citations
- SpatiaLab: Can Vision-Language Models Perform Spatial Reasoning in the Wild?Azmine Toushik Wasi, Wahid Faisal, Abdur Rahman, Mahfuz Ahmed Anik et al.ICLR 2026 · 13 citations
- SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous DrivingPeizheng Li, Zhenghao Zhang, David Holtz, Hang Yu et al.CVPR 2026 · 32 citations
