Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?
Pingyue Zhang, Zihan Huang, Yue Wang, Jieyu Zhang, Letian Xue, Zihan Wang, Qineng Wang, Keshigeyan Chandrasegaran, Ruohan Zhang, Yejin Choi, Ranjay Krishna, Jiajun Wu
Abstract
Spatial embodied intelligence often operates under partial observability, where agents must act to acquire missing information rather than passively consume complete observations. In such settings, progress depends on actively selecting informative actions that reduce uncertainty and support the construction of spatial understanding. While multimodal foundation models have shown strong performance on passive multimodal perception and reasoning tasks, their ability to support active, self-directed exploration under partial observability has not been systematically studied. In particular, it remains unclear whether and how these models can decide what to observe next in order to build and maintain a coherent spatial belief over time. We therefore propose THEORY OF SPACE, defined as an agent's ability to actively acquire information through self-directed, active exploration and to construct, revise, and exploit a spatial belief from sequential, partial observations. We implement THEORY OF SPACE using a benchmark with textual and visual environments. Rather than solving specific tasks, the goal is curiositydriven exploration to build a complete, accurate spatial belief. A core innovation is spatial belief probing: we prompt it to reveal its internal spatial belief as a cognitive map at each step, letting us measure the quality of its underlying spatial model. Our evaluation of state-of-the-art models on a suite of downstream tasks reveals critical bottlenecks: (1) The Active-Passive Gap: Performance degrades when agents must autonomously gather information (e.g., GPT-5.2: 0.57→0.46); (2) Inefficiency: Models explore in an unsystematic way and with high redundancy, failing to match the efficiency of program-based proxies while producing no better results. Through belief probing, we diagnose that perception acts as an initial bottleneck, yet global beliefs suffer further from instability that causes spatial knowledge to degrade over time. Finally, using a false belief paradigm to test belief revision, we uncover Belief Inertia where agents fail to overwrite obsolete priors. This issue exists in text agents but is notably severe in vision-based models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2456510e-d13f-4b48-8356-1b3e61844cbcBuilds on21
- SpatialRGPT: Grounded Spatial Reasoning in Vision-Language ModelsAn-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo et al.NeurIPS 2024 · 412 citations
- TEACh: Task-Driven Embodied Agents That ChatAishwarya Padmakumar, Jesse Thomason, Ayush Shrivastava, Patrick Lange et al.AAAI 2022 · 251 citations
- Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingAlexander Ku, Peter Anderson, Roma Patel, Eugene Ie et al.EMNLP 2020 · 208 citations
- MMSI-Bench: A Benchmark for Multi-Image Spatial IntelligenceSihan Yang, Runsen Xu, Yiman Xie, Sizhe Yang et al.ICLR 2026 · 195 citations
- StepGame: A New Benchmark for Robust Multi-Hop Spatial Reasoning in TextsZhengxiang Shi, Qiang Zhang, Aldo LipaniAAAI 2022 · 100 citations
Related papers
- ProcWorld: Benchmarking Large Model Planning in Reachability-Constrained EnvironmentsDong Wang, Xinghang Li, Zhengshen Zhang, Jirong Liu et al.EMNLP 2025
- EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive HierarchyJinzhao Li, Yinuo Chen, Dongxu Piao, Panwang Pan et al.CVPR 2026 · 2 citations
- BEAR: Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and DiagnosisYu Qi, Haibo Zhao, Ziyu Guo, Siyuan Ma et al.ICML 2026
- Towards Physics-informed Spatial Intelligence with Human Priors: An Autonomous Driving Pilot StudyGuanlin Wu, Boyan Su, Yang Zhao, Pu Wang et al.NeurIPS 2025 · 2 citations
- Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language ModelsJiayu Wang, Yifei Ming, Zhenmei Shi, Vibhav Vineet et al.NeurIPS 2024 · 166 citations
