Enhancing Spatial Reasoning Through Visual and Textual Thinking
Xun Liang, Xin Guo, Zhongming Jin, Weihang Pan, Penghui Shang, Deng Cai, Binbin Lin, Jieping Ye
摘要
The spatial reasoning task aims to reason about the spatial relationships in 2D and 3D space, which is a fundamental capability for Visual Question Answering (VQA) and robotics. Although vision language models (VLMs) have developed rapidly in recent years, they are still struggling with the spatial reasoning task. In this paper, we introduce a method that can enhance Spatial reasoning through Visual and Textual thinking Simultaneously (SpatialVTS). In the spatial visual thinking phase, our model is trained to generate locationrelated specific tokens of essential targets automatically. Not only are the objects mentioned in the problem addressed, but also the potential objects related to the reasoning are considered. During the spatial textual thinking phase, our model conducts long-term thinking based on visual cues and dialogues, gradually inferring the answers to spatial reasoning problems. To effectively support the model's training, we perform manual corrections to the existing spatial reasoning dataset, eliminating numerous incorrect labels resulting from automatic annotation, restructuring the data input format to enhance generalization ability, and developing thinking processes with logical reasoning details. Without introducing additional information (such as masks or depth), our model's overall average level in several spatial understanding tasks has significantly improved compared with other models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- VisMem: Latent Vision Memory Unlocks Potential of Vision-Language ModelsXinlei Yu, Chengming Xu, Guibin Zhang, Zhangquan Chen 等CVPR 2026 · 被引用 30 次
- Self-Prophetic Decoding to Unlock Visual Search in LVLMsZhendong He, Qiyuan Dai, Guanbin Li, Liang Lin 等ICML 2026
它引用的顶会 Paper11
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- Visual-RFT: Visual Reinforcement Fine-TuningZiyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong 等ICCV 2025 · 被引用 563 次
- SpatialRGPT: Grounded Spatial Reasoning in Vision-Language ModelsAn-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo 等NeurIPS 2024 · 被引用 412 次
- Osprey: Pixel Understanding with Visual Instruction TuningYuqian Yuan, Wentong Li, Jian Liu, Dongqi Tang 等CVPR 2024 · 被引用 39 次
- LlaVA-CoT: Let Vision Language Models Reason Step-By-StepGuowei Xu, Peng Jin, Ziang Wu, Hao Li 等ICCV 2025 · 被引用 37 次
相关 Paper
- SpatiaLQA: A Benchmark for Evaluating Spatial Logical Reasoning in Vision-Language ModelsYuechen Xie, Xiaoyan Zhang, Yicheng Shan, Zhu Hao 等CVPR 2026 · 被引用 8 次
- SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D PriorsChenyang Ma, Kai Lu, Ta Ying Cheng, Niki Trigoni 等NeurIPS 2024 · 被引用 82 次
- SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic DataMichael Ogezi, Freda ShiACL 2025 · 被引用 19 次
- Spatial Understanding from Videos: Structured Prompts Meet Simulation DataHaoyu Zhang, Meng Liu, Zaijing Li, Haokun Wen 等NeurIPS 2025 · 被引用 31 次
- SpatialReasoner: Towards Explicit and Generalizable 3D Spatial ReasoningWufei Ma, Yu-Cheng Chou, Qihao Liu, Xingrui Wang 等NeurIPS 2025 · 被引用 77 次
