What's "up" with vision-language models? Investigating their struggle with spatial reasoning
Amita Kamath, Jack Hessel, Kai-Wei Chang
摘要
Recent vision-language (VL) models are powerful, but can they reliably distinguish "right" from "left"? We curate three new corpora to quantify model comprehension of such basic spatial relations. These tests isolate spatial reasoning more precisely than existing datasets like VQAv2, e.g., our What'sUp benchmark contains sets of photographs varying only the spatial relations of objects, keeping their identity fixed (see Figure 1 : models must comprehend not only the usual case of a dog under a table, but also, the same dog on top of the same table). We evaluate 18 VL models, finding that all perform poorly, e.g., BLIP finetuned on VQAv2, which nears human parity on VQAv2, achieves 56% accuracy on our benchmarks vs. humans at 99%. We conclude by studying causes of this surprising behavior, finding: 1) that popular vision-language pretraining corpora like LAION-2B contain little reliable data for learning spatial relationships; and 2) that basic modeling interventions like up-weighting preposition-containing instances or fine-tuning on our corpora are not sufficient to address the challenges our benchmarks pose. We are hopeful that these corpora will facilitate further research, and we release our data and code at https://github.com/amitakamath/ whatsup_vlms .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper105
- RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for RoboticsEnshen Zhou, Jingkun An, Cheng Chi, Yi Han 等NeurIPS 2025 · 被引用 159 次
- Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual DrawingJunfei Wu, Jian Guan, Kaituo Feng, Qiang Liu 等NeurIPS 2025 · 被引用 153 次
- OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language ModelsMengdi Jia, Zekun Qi, Shaochen Zhang, Wenyao Zhang 等ICLR 2026 · 被引用 109 次
- SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language ModelsHongxing Li, Dingming Li, Zixuan Wang, Yuchen Yan 等ICLR 2026 · 被引用 75 次
- Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View ScenesMohsen Gholami, Ahmad Rezaei, Zhou Weimin, Sitong Mao 等ICLR 2026 · 被引用 67 次
它引用的顶会 Paper11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- FLAVA: A Foundational Language And Vision Alignment ModelAmanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon 等CVPR 2022 · 被引用 483 次
- TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question AnsweringYushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang 等ICCV 2023 · 被引用 400 次
相关 Paper
- InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language ModelsNianchen Deng, Lixin Gu, Shenglong Ye, Yinan He 等ICLR 2026 · 被引用 32 次
- SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic DataMichael Ogezi, Freda ShiACL 2025 · 被引用 19 次
- Expand VSR Benchmark for VLLM to Expertize in Spatial RulesPeijin Xie, Lin Sun, Bingquan Liu, Dexin Wang 等AAAI 2025
- SPHERE: Unveiling Spatial Blind Spots in Vision-Language Models Through Hierarchical EvaluationWenyu Zhang, Wei En Ng, Lixin Ma, Yuwen Wang 等ACL 2025 · 被引用 20 次
- SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMsYuyou Zhang, Radu Corcodel, Chiori Hori, Anoop Cherian 等ICLR 2026 · 被引用 11 次
