SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data
Michael Ogezi, Freda Shi
摘要
Vision-language models (VLMs) work well in tasks ranging from image captioning to visual question answering (VQA), yet they struggle with spatial reasoning, a key skill for understanding our physical world that humans excel at. We find that spatial relations are generally rare in widely used VL datasets, with only a few being well represented while most form a long tail of underrepresented relations. This gap leaves VLMs ill-equipped to handle diverse spatial relationships. To bridge it, we construct a synthetic VQA dataset focused on spatial reasoning generated from hyperdetailed image descriptions in Localized Narratives, DOCCI, and PixMo-Cap. Our dataset consists of 455k samples containing 3.4 million QA pairs. Trained on this dataset, our Spatial-Reasoning Enhanced (SpaRE) VLMs show strong improvements on spatial reasoning benchmarks, achieving up to a 49% performance gain on the What's Up benchmark, while maintaining strong results on general tasks. Our work narrows the gap between human and VLM spatial reasoning and makes VLMs more capable in real-world tasks such as robotics and navigation. We look to share our code and dataset in due course.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Reason, Then Re-reason: Cross-view Revisiting Improves Spatial ReasoningChaofan Ma, Zhenjie Mao, Yuhuan Yang, Fanqin Zeng 等ICML 2026 · 被引用 1 次
- Keep it SymPL: Symbolic Projective Layout for Allocentric Spatial Reasoning in Vision-Language ModelsJaeyun Jang, Seunghui Shin, Taeho Park, Hyoseok HwangCVPR 2026 · 被引用 1 次
- Closing the Spatial Execution Gap in Digital Whiteboards via Verifiable Reinforcement LearningChang Liu, Benjamin Wagley, Zibo Wang, Mehmet E. Belviranli 等ACL 2026
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesDouwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami 等NeurIPS 2020 · 被引用 1,022 次
- Objects365: A Large-Scale, High-Quality Dataset for Object DetectionShuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng 等ICCV 2019 · 被引用 1,018 次
相关 Paper
- Enhancing Spatial Reasoning Through Visual and Textual ThinkingXun Liang, Xin Guo, Zhongming Jin, Weihang Pan 等AAAI 2026
- Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language ModelsXinmiao Huang, Qisong He, Zhenglin Huang, Boxuan Wang 等ICLR 2026 · 被引用 8 次
- TopViewRS: Vision-Language Models as Top-View Spatial ReasonersChengzu Li, Caiqi Zhang, Han Zhou, Nigel Collier 等EMNLP 2024 · 被引用 5 次
- Learning Multi-View Spatial Reasoning from Cross-View RelationsSuchae Jeong, Jaehwi Song, Haeone Lee, Hanna Kim 等CVPR 2026
- An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal ModelsFatemeh Shiri, Xiao-Yu Guo, Mona Far, Xin Yu 等EMNLP 2024 · 被引用 7 次
