Advancing Spatial Reasoning in Large Language Models: An In-Depth Evaluation and Enhancement Using the StepGame Benchmark
Fangjun Li, David C. Hogg, Anthony G. Cohn
Abstract
Artificial intelligence (AI) has made remarkable progress across various domains, with large language models like ChatGPT gaining substantial attention for their human-like text-generation capabilities. Despite these achievements, spatial reasoning remains a significant challenge for these models. Benchmarks like StepGame evaluate AI spatial reasoning, where ChatGPT has shown unsatisfactory performance. However, the presence of template errors in the benchmark has an impact on the evaluation results. Thus there is potential for ChatGPT to perform better if these template errors are addressed, leading to more accurate assessments of its spatial reasoning capabilities. In this study, we refine the StepGame benchmark, providing a more accurate dataset for model evaluation. We analyze GPT's spatial reasoning performance on the rectified benchmark, identifying proficiency in mapping natural language text to spatial relations but limitations in multi-hop reasoning. We provide a flawless solution to the benchmark by combining template-to-relation mapping with logic-based reasoning. This combination demonstrates proficiency in performing qualitative reasoning on StepGame without encountering any errors. We then address the limitations of GPT models in spatial reasoning. We deploy Chainof-thought and Tree-of-thoughts prompting strategies, offering insights into GPT's "cognitive process", and achieving remarkable improvements in accuracy. Our investigation not only sheds light on model deficiencies but also proposes enhancements, contributing to the advancement of AI with more robust spatial reasoning capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3c868fb4-9fd2-4d3f-9264-314358928275Cited by top-tier papers4
- PRIMT: Preference-based Reinforcement Learning with Multimodal Feedback and Trajectory Synthesis from Foundation ModelsRuiqi Wang, Dezhong Zhao, Ziqin Yuan, Tianyu Shao et al.NeurIPS 2025 · 9 citations
- Seeing What Matters: A Training-Free Self-Guided Framework for Multimodal Detail Perception and ReasoningMingjie Ma, yichao ma, Zhong Yang, Guohui LiCVPR 2026
- Large Language and Reasoning Models are Shallow Disjunctive ReasonersIrtaza Khalid, Amir Masoud Nourollah, Steven SchockaertACL 2025
- How Do LLMs and VLMs Understand Viewpoint Rotation Without Vision? An Interpretability StudyZhen Yang, Ping Jian, Zhongbin Guo, Zuming Zhang et al.ACL 2026
Builds on8
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Least-to-Most Prompting Enables Complex Reasoning in Large Language ModelsDenny Zhou, Nathanael Schärli, Le Hou, Jason Wei et al.ICLR 2023 · 318 citations
- Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language ModelsLei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu et al.ACL 2023 · 249 citations
- Automatic Chain of Thought Prompting in Large Language ModelsZhuosheng Zhang, Aston Zhang, Mu Li, Alex SmolaICLR 2023 · 234 citations
Related papers
- FoREST: Frame of Reference Evaluation in Spatial Reasoning TasksTanawan Premsri, Parisa KordjamshidiEMNLP 2025
- MMSI-Bench: A Benchmark for Multi-Image Spatial IntelligenceSihan Yang, Runsen Xu, Yiman Xie, Sizhe Yang et al.ICLR 2026 · 195 citations
- GeoGramBench: Benchmarking the Geometric Program Reasoning in Modern LLMsShixian Luo, Zhu zezhou, Yu Yuan, Yuncheng Yang et al.ICLR 2026 · 15 citations
- SpaRC and SpaRP: Spatial Reasoning Characterization and Path Generation for Understanding Spatial Reasoning Capability of Large Language ModelsMd Imbesat Hassan Rizvi, Xiaodan Zhu, Iryna GurevychACL 2024
- One Cognitive Loop Is Enough: SODA unlocks Pure-Text Spatial Reasoning in Large Language ModelsShunwen Bai, Jiahuan Zhang, Haoran Huang, Yurun Wang et al.ACL 2026
