An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models
Fatemeh Shiri, Xiao-Yu Guo, Mona Far, Xin Yu, Reza Haf, Yuan-Fang Li
Abstract
Large Multimodal Models (LMMs) have achieved strong performance across a range of vision and language tasks.However, their spatial reasoning capabilities are underinvestigated.In this paper, we construct a novel VQA dataset, Spatial-MM, to comprehensively study LMMs' spatial understanding and reasoning capabilities.Our analyses on object-relationship and multi-hop reasoning reveal several important findings.Firstly, bounding boxes and scene graphs, even synthetic ones, can significantly enhance LMMs' spatial reasoning.Secondly, LMMs struggle more with questions posed from the human perspective than the camera perspective about the image.Thirdly, chain of thought (CoT) prompting does not improve model performance on complex multi-hop questions involving spatial relations.Lastly, our perturbation analysis on GQA-spatial reveals that LMMs are much stronger at basic object detection than complex spatial reasoning.We believe our new benchmark dataset and in-depth analyses can spark further research on LMMs spatial reasoning. 11 Spatial-MM benchmark is available at: https://github. com/FatemehShiri/Spatial-MM Where is the bicycle from the woman's perspective?A. Front B. Behind C.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f3b66fb1-edc2-4afd-8651-a893df492ff4Cited by top-tier papers36
- MMSI-Bench: A Benchmark for Multi-Image Spatial IntelligenceSihan Yang, Runsen Xu, Yiman Xie, Sizhe Yang et al.ICLR 2026 · 195 citations
- RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for RoboticsEnshen Zhou, Jingkun An, Cheng Chi, Yi Han et al.NeurIPS 2025 · 159 citations
- OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language ModelsMengdi Jia, Zekun Qi, Shaochen Zhang, Wenyao Zhang et al.ICLR 2026 · 109 citations
- Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language ModelsRunsen Xu, Weiyao Wang, Hao Tang, Xingyu Chen et al.CVPR 2026 · 64 citations
- SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language ModelsPingyi Chen, Yujing Lou, Shen Cao, Jinhui Guo et al.NeurIPS 2025 · 24 citations
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Grounding Multimodal Large Language Models to the WorldZhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao et al.ICLR 2024 · 1,170 citations
- Least-to-Most Prompting Enables Complex Reasoning in Large Language ModelsDenny Zhou, Nathanael Schärli, Le Hou, Jason Wei et al.ICLR 2023 · 318 citations
Related papers
- Can Multimodal Large Language Models Understand Spatial Relations?Jingping Liu, Ziyan Liu, Zhedong Cen, Yan Zhou et al.ACL 2025 · 16 citations
- TopViewRS: Vision-Language Models as Top-View Spatial ReasonersChengzu Li, Caiqi Zhang, Han Zhou, Nigel Collier et al.EMNLP 2024 · 5 citations
- Grounded Chain-of-Thought for Multimodal Large Language ModelsQiong Wu, Xiangcong Yang, Yiyi Zhou, Chenxin Fang et al.CVPR 2026 · 55 citations
- SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal ModelsWufei Ma, Luoxin Ye, Celso M. de Melo, Alan L. Yuille et al.CVPR 2025
- 3DSRBENCH: A Comprehensive 3D Spatial Reasoning BenchmarkWufei Ma, Haoyu Chen, Guofeng Zhang, Yu-Cheng Chou et al.ICCV 2025 · 15 citations
