Advancing Spatial Reasoning in Large Language Models: An In-Depth Evaluation and Enhancement Using the StepGame Benchmark
Fangjun Li, David C. Hogg, Anthony G. Cohn
摘要
Artificial intelligence (AI) has made remarkable progress across various domains, with large language models like ChatGPT gaining substantial attention for their human-like text-generation capabilities. Despite these achievements, spatial reasoning remains a significant challenge for these models. Benchmarks like StepGame evaluate AI spatial reasoning, where ChatGPT has shown unsatisfactory performance. However, the presence of template errors in the benchmark has an impact on the evaluation results. Thus there is potential for ChatGPT to perform better if these template errors are addressed, leading to more accurate assessments of its spatial reasoning capabilities. In this study, we refine the StepGame benchmark, providing a more accurate dataset for model evaluation. We analyze GPT's spatial reasoning performance on the rectified benchmark, identifying proficiency in mapping natural language text to spatial relations but limitations in multi-hop reasoning. We provide a flawless solution to the benchmark by combining template-to-relation mapping with logic-based reasoning. This combination demonstrates proficiency in performing qualitative reasoning on StepGame without encountering any errors. We then address the limitations of GPT models in spatial reasoning. We deploy Chainof-thought and Tree-of-thoughts prompting strategies, offering insights into GPT's "cognitive process", and achieving remarkable improvements in accuracy. Our investigation not only sheds light on model deficiencies but also proposes enhancements, contributing to the advancement of AI with more robust spatial reasoning capabilities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- PRIMT: Preference-based Reinforcement Learning with Multimodal Feedback and Trajectory Synthesis from Foundation ModelsRuiqi Wang, Dezhong Zhao, Ziqin Yuan, Tianyu Shao 等NeurIPS 2025 · 被引用 9 次
- Seeing What Matters: A Training-Free Self-Guided Framework for Multimodal Detail Perception and ReasoningMingjie Ma, yichao ma, Zhong Yang, Guohui LiCVPR 2026
- Large Language and Reasoning Models are Shallow Disjunctive ReasonersIrtaza Khalid, Amir Masoud Nourollah, Steven SchockaertACL 2025
- How Do LLMs and VLMs Understand Viewpoint Rotation Without Vision? An Interpretability StudyZhen Yang, Ping Jian, Zhongbin Guo, Zuming Zhang 等ACL 2026
它引用的顶会 Paper8
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Least-to-Most Prompting Enables Complex Reasoning in Large Language ModelsDenny Zhou, Nathanael Schärli, Le Hou, Jason Wei 等ICLR 2023 · 被引用 318 次
- Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language ModelsLei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu 等ACL 2023 · 被引用 249 次
- Automatic Chain of Thought Prompting in Large Language ModelsZhuosheng Zhang, Aston Zhang, Mu Li, Alex SmolaICLR 2023 · 被引用 234 次
相关 Paper
- FoREST: Frame of Reference Evaluation in Spatial Reasoning TasksTanawan Premsri, Parisa KordjamshidiEMNLP 2025
- MMSI-Bench: A Benchmark for Multi-Image Spatial IntelligenceSihan Yang, Runsen Xu, Yiman Xie, Sizhe Yang 等ICLR 2026 · 被引用 195 次
- GeoGramBench: Benchmarking the Geometric Program Reasoning in Modern LLMsShixian Luo, Zhu zezhou, Yu Yuan, Yuncheng Yang 等ICLR 2026 · 被引用 15 次
- SpaRC and SpaRP: Spatial Reasoning Characterization and Path Generation for Understanding Spatial Reasoning Capability of Large Language ModelsMd Imbesat Hassan Rizvi, Xiaodan Zhu, Iryna GurevychACL 2024
- One Cognitive Loop Is Enough: SODA unlocks Pure-Text Spatial Reasoning in Large Language ModelsShunwen Bai, Jiahuan Zhang, Haoran Huang, Yurun Wang 等ACL 2026
