Weakly Supervised Relative Spatial Reasoning for Visual Question Answering
Pratyay Banerjee, Tejas Gokhale, Yezhou Yang, Chitta Baral
Abstract
Vision-and-language (V&L) reasoning necessitates perception of visual concepts such as objects and actions, understanding semantics and language grounding, and reasoning about the interplay between the two modalities. One crucial aspect of visual reasoning is spatial understanding, which involves understanding relative locations of objects, i.e. implicitly learning the geometry of the scene. In this work, we evaluate the faithfulness of V&L models to such geometric understanding, by formulating the prediction of pair-wise relative locations of objects as a classification as well as a regression task. Our findings suggest that state-of-the-art transformer-based V&L models lack sufficient abilities to excel at this task. Motivated by this, we design two objectives as proxies for 3D spatial reasoning (SR) – object centroid estimation, and relative position estimation, and train V&L with weak supervision from off-the-shelf depth estimators. This leads to considerable improvements in accuracy for the "GQA" visual question answering challenge (in fully supervised, few-shot, and O.O.D settings) as well as improvements in relative spatial reasoning. Code and data will be released here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a9d1e4a9-dc31-4da4-84bd-6fa5c08a9942Cited by top-tier papers4
- Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMsKanchana Ranasinghe, Satya Narayan Shukla, Omid Poursaeed, Michael S. Ryoo et al.CVPR 2024 · 21 citations
- Motion-R1: Enhancing Motion Generation with Decomposed Chain-of-Thought and RL BindingRunqi Ouyang, Haoyun Li, Zhenyuan Zhang, Xiaofeng Wang et al.ICLR 2026 · 6 citations
- TopViewRS: Vision-Language Models as Top-View Spatial ReasonersChengzu Li, Caiqi Zhang, Han Zhou, Nigel Collier et al.EMNLP 2024 · 5 citations
- Logical Implications for Visual Question Answering ConsistencySergio Tascon-Morales, Pablo Márquez-Neila, Raphael SznitmanCVPR 2023
Builds on1
Related papers
- G^2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial ReasoningWenbo hu, JINGLI LIN, Yilin Long, Yunlong Ran et al.CVPR 2026
- SpatiaLQA: A Benchmark for Evaluating Spatial Logical Reasoning in Vision-Language ModelsYuechen Xie, Xiaoyan Zhang, Yicheng Shan, Zhu Hao et al.CVPR 2026 · 8 citations
- Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric ReasoningChun-Hsiao Yeh, Shengyi Qian, Manchen Wang, Yi Ma et al.CVPR 2026 · 1 citation
- Proxy3D: Efficient 3D Representations for Vision-Language Models via Semantic Clustering and AlignmentJerry Jiang, Haowen Sun, Denis A. Gudovskiy, Yohei Nakata et al.CVPR 2026 · 3 citations
- An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal ModelsFatemeh Shiri, Xiao-Yu Guo, Mona Far, Xin Yu et al.EMNLP 2024 · 7 citations
