MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
Sihan Yang, Runsen Xu, Yiman Xie, Sizhe Yang, Mo Li, Jingli Lin, Chenming Zhu, Xiaochen Chen, Haodong Duan, Xiangyu Yue, Dahua Lin, Tai Wang, Jiangmiao Pang
摘要
Spatial intelligence is essential for multimodal large language models (MLLMs) operating in the complex physical world. Existing benchmarks, however, probe only single-image relations and thus fail to assess the multi-image spatial reasoning that real-world deployments demand. We introduce MMSI-Bench, a VQA benchmark dedicated to multi-image spatial intelligence. Six 3D-vision researchers spent more than 300 hours meticulously crafting 1,000 challenging, unambiguous multiple-choice questions from over 120,000 images, each paired with carefully designed distractors and a stepwise reasoning process. We conduct extensive experiments and evaluate 37 open-source and proprietary MLLMs, observing a wide gap: the strongest open-source model attains roughly 30% accuracy and OpenAI's GPT-5 reasoning model reaches 40%, while humans score 97%. These results underscore the challenging nature of MMSI-Bench and the substantial headroom for future research. Leveraging the annotated reasoning processes, we also provide an automated error analysis pipeline that diagnoses four dominant failure modes, including (1) grounding errors, (2) overlap-matching and scene-reconstruction errors, (3) situation-transformation reasoning errors, and (4) spatial-logic errors, offering insights for advancing spatial intelligence.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper48
- Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual DrawingJunfei Wu, Jian Guan, Kaituo Feng, Qiang Liu 等NeurIPS 2025 · 被引用 153 次
- Scaling Spatial Intelligence with Multimodal Foundation ModelsZhongang Cai, Wang Ruisi, Chenyang Gu, Fanyi Pu 等CVPR 2026 · 被引用 81 次
- Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View ScenesMohsen Gholami, Ahmad Rezaei, Zhou Weimin, Sitong Mao 等ICLR 2026 · 被引用 67 次
- Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited ViewsZhangquan Chen, Manyuan Zhang, Xinlei Yu, Xufang Luo 等CVPR 2026 · 被引用 61 次
- Geometrically-Constrained Agent for Spatial ReasoningZeren Chen, Xiaoya Lu, Zhijie Zheng, Pengrui Li 等CVPR 2026 · 被引用 29 次
它引用的顶会 Paper23
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsPeter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo 等NeurIPS 2024 · 被引用 1,004 次
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis 等CVPR 2022 · 被引用 525 次
- SpatialRGPT: Grounded Spatial Reasoning in Vision-Language ModelsAn-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo 等NeurIPS 2024 · 被引用 412 次
- Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial IntelligenceDiankun Wu, Fangfu Liu, Yi-Hsin Hung, Yueqi DuanNeurIPS 2025 · 被引用 245 次
相关 Paper
- 3DSRBENCH: A Comprehensive 3D Spatial Reasoning BenchmarkWufei Ma, Haoyu Chen, Guofeng Zhang, Yu-Cheng Chou 等ICCV 2025 · 被引用 15 次
- Can Multimodal Large Language Models Understand Spatial Relations?Jingping Liu, Ziyan Liu, Zhedong Cen, Yan Zhou 等ACL 2025 · 被引用 16 次
- SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMsSiting Wang, Minnan Pei, Luoyang Sun, Cheng Deng 等ICLR 2026 · 被引用 8 次
- Thinking in Structures: Evaluating Spatial Intelligence in Constraint-Governed SpacesChen Yang, Guanxin Lin, Youquan He, Peiyao Chen 等ICML 2026
- From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMsMingrui Wu, Zhaozhi Wang, Fangjinhua Wang, Jiaolong Yang 等CVPR 2026 · 被引用 11 次
