SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence
Ziyang Gong, Wenhao Li, Xianzheng Ma, Songyuan Li, Zhaokai Wang, Songze Li, Jiayi Ji, Xue Yang, Gen Luo, Junchi Yan, Rongrong Ji
Abstract
Multimodal Large Language Models (MLLMs) have achieved remarkable progress in various multimodal tasks. To pursue higher intelligence in space, MLLMs require integrating multiple atomic spatial capabilities to handle complex and dynamic tasks. However, existing benchmarks struggle to comprehensively evaluate the spatial intelligence of common MLLMs from the atomic level to the compositional level. To fill this gap, we present SpaCE-10, a comprehensive benchmark for compositional spatial evaluations. In SpaCE-10, we define 10 atomic spatial capabilities, which are combined to form 8 compositional capabilities. Based on these definitions, we propose a novel hierarchical annotation pipeline to generate high-quality and diverse question-answer (QA) pairs. With over 150+ hours of human expert effort, we obtain over 5k QA pairs for 811 real indoor scenes in SpaCE-10, which covers various evaluation settings like point cloud input and multi-choice QA. We conduct an extensive evaluation of common MLLMs on SpaCE-10 and find that even the most advanced MLLM still lags behind humans by large margins. Through our careful study, we also draw several significant findings that benefit the MLLM community. For example, we reveal that the shortcoming of counting capability greatly limits the compositional spatial capabilities of existing MLLMs. We will release the code and benchmark soon.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 44944349-3b66-448d-b001-f5e648c38bafCited by top-tier papers8
- Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image ModelsZengbin Wang, Xuecai Hu, Yong Wang, Feng Xiong et al.ICLR 2026 · 13 citations
- SpatialScore: Towards Comprehensive Evaluation for Spatial IntelligenceHaoning Wu, Xiao Huang, Yaohui Chen, Ya Zhang et al.CVPR 2026 · 12 citations
- From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMsMingrui Wu, Zhaozhi Wang, Fangjinhua Wang, Jiaolong Yang et al.CVPR 2026 · 11 citations
- SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMsYuyou Zhang, Radu Corcodel, Chiori Hori, Anoop Cherian et al.ICLR 2026 · 11 citations
- CrossCheck-Bench: Diagnosing Compositional Failures in Multimodal Conflict ResolutionBaoliang Tian, Yuxuan Si, Jilong Wang, Lingyao Li et al.AAAI 2026 · 2 citations
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsPeter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo et al.NeurIPS 2024 · 1,004 citations
- 3D-LLM: Injecting the 3D World into Large Language ModelsYining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng et al.NeurIPS 2023 · 662 citations
- SpatialRGPT: Grounded Spatial Reasoning in Vision-Language ModelsAn-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo et al.NeurIPS 2024 · 412 citations
Related papers
- MMSI-Bench: A Benchmark for Multi-Image Spatial IntelligenceSihan Yang, Runsen Xu, Yiman Xie, Sizhe Yang et al.ICLR 2026 · 195 citations
- SpaCE-Eval: A Benchmark for Real-World Multi-Modal ReasoningXuyou Yang, Yucheng Zhao, Wenxuan Zhang, Immanuel KohICLR 2026
- VSP: Diagnosing the Dual Challenges of Perception and Reasoning in Spatial Planning Tasks for MLLMSQiucheng Wu, Handong Zhao, Michael Saxon, Trung Bui et al.ICCV 2025 · 1 citation
- Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language ModelsJiayu Wang, Yifei Ming, Zhenmei Shi, Vibhav Vineet et al.NeurIPS 2024 · 166 citations
- Can Multimodal Large Language Models Understand Spatial Relations?Jingping Liu, Ziyan Liu, Zhedong Cen, Yan Zhou et al.ACL 2025 · 16 citations
