FoREST: Frame of Reference Evaluation in Spatial Reasoning Tasks
Tanawan Premsri, Parisa Kordjamshidi
Abstract
Spatial reasoning is a fundamental aspect of human intelligence. One key concept in spatial cognition is the Frame of Reference (FoR), which identifies the perspective of spatial expressions. Despite its significance, FoR has received limited attention in AI models that need spatial intelligence. There is a lack of dedicated benchmarks and in-depth evaluation of large language models (LLMs) in this area. To address this issue, we introduce the Frame of Reference Evaluation in Spatial Reasoning Tasks (FoREST) benchmark, designed to assess FoR comprehension in LLMs. We evaluate LLMs on answering questions that require FoR comprehension and layout generation in textto-image models using FoREST. Our results reveal a notable performance gap across different FoR classes in various LLMs, affecting their ability to generate accurate layouts for text-toimage generation. This highlights critical shortcomings in FoR comprehension. To improve FoR understanding, we propose Spatial-Guided prompting, which improves LLMs' ability to extract primitive spatial concepts and relations. Our proposed method improves overall performance across spatial reasoning tasks. Context Generation List of Objects Locatum (L) Relatum (R) <L> <relation> <R> A cat is to the right of a dog from the dog's perspective. A dog is facing toward the camera. Q: Based on camera angle, where is the cat from the dog's position? A: Left Q: In the dog view, how is the cat positioned in relation to the dog? A: Right Camera's perspective Relatum's perspective Visualization A cat is to the right of a dog.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d45ce5e2-cfc4-4f31-a020-6f4fe572a918Cited by top-tier papers1
Ask how each one uses itBuilds on7
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- StepGame: A New Benchmark for Robust Multi-Hop Spatial Reasoning in TextsZhengxiang Shi, Qiang Zhang, Aldo LipaniAAAI 2022 · 100 citations
- Transfer Learning with Synthetic Corpora for Spatial Role Labeling and ReasoningRoshanak Mirzaee, Parisa KordjamshidiEMNLP 2022 · 10 citations
- SpaRC and SpaRP: Spatial Reasoning Characterization and Path Generation for Understanding Spatial Reasoning Capability of Large Language ModelsMd Imbesat Hassan Rizvi, Xiaodan Zhu, Iryna GurevychACL 2024
Related papers
- FloorplanQA: A Benchmark for Spatial Reasoning in LLMs using Structured RepresentationsFedor Rodionov, Abdelrahman Eldesokey, Michael Birsak, John Femiani et al.ICML 2026 · 14 citations
- Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language ModelsJiayu Wang, Yifei Ming, Zhenmei Shi, Vibhav Vineet et al.NeurIPS 2024 · 166 citations
- NoReGeo: Non-Reasoning Geometry BenchmarkIrina Abdullaeva, Anton Vasiliuk, Elizaveta Goncharova, Temurbek Rahmatullaev et al.AAAI 2026
- Towards Physics-informed Spatial Intelligence with Human Priors: An Autonomous Driving Pilot StudyGuanlin Wu, Boyan Su, Yang Zhao, Pu Wang et al.NeurIPS 2025 · 2 citations
- Advancing Spatial Reasoning in Large Language Models: An In-Depth Evaluation and Enhancement Using the StepGame BenchmarkFangjun Li, David C. Hogg, Anthony G. CohnAAAI 2024 · 60 citations
