Evaluating the World Model Implicit in a Generative Model
Keyon Vafa, Justin Y. Chen, Ashesh Rambachan, Jon M. Kleinberg, Sendhil Mullainathan
Abstract
Recent work suggests that large language models may implicitly learn world models. How should we assess this possibility? We formalize this question for the case where the underlying reality is governed by a deterministic finite automaton. This includes problems as diverse as simple logical reasoning, geographic navigation, game-playing, and chemistry. We propose new evaluation metrics for world model recovery inspired by the classic Myhill-Nerode theorem from language theory. We illustrate their utility in three domains: game playing, logic puzzles, and navigation. In all domains, the generative models we consider do well on existing diagnostics for assessing world models, but our evaluation metrics reveal their world models to be far less coherent than they appear. Such incoherence creates fragility: using a generative model to solve related but subtly different tasks can lead to failures. Building generative models that meaningfully capture the underlying logic of the domains they model would be immensely valuable; our results suggest new ways to assess how close a given model is to that goal.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd6db38b-e483-4063-b5be-445456b02b90Cited by top-tier papers22
- One Life to Learn: Inferring Symbolic World Models for Stochastic Environments from Unguided ExplorationZaid Khan, Archiki Prasad, Elias Stengel-Eskin, Jaemin Cho et al.ICLR 2026 · 12 citations
- Finite State Automata Inside Transformers with Chain-of-Thought: A Mechanistic Study on State TrackingYifan Zhang, Wenyu Du, Dongming Jin, Jie Fu et al.ACL 2025 · 11 citations
- Characterizing the Effect of Noise in Language Generation in the LimitAaron Li, Ian ZhangICML 2026 · 4 citations
- Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated SurveyIvan Vegner, Sydelle de Souza, Valentin Forch, Martha Lewis et al.ACL 2025 · 3 citations
- Language Models Struggle to Use Representations Learned In-ContextMichael A. Lepori, Tal Linzen, Ann Yuan, Katja FilippovaACL 2026 · 3 citations
Builds on12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Leveraging Pre-trained Large Language Models to Construct and Utilize World Models for Model-based Task PlanningLin Guan, Karthik Valmeekam, Sarath Sreedharan, Subbarao KambhampatiNeurIPS 2023 · 347 citations
- Mapping Language Models to Grounded Conceptual SpacesRoma Patel, Ellie PavlickICLR 2022 · 197 citations
- The Illusion of State in State-Space ModelsWilliam Merrill, Jackson Petty, Ashish SabharwalICML 2024 · 157 citations
Related papers
- DAG-Math: Graph-of-Thought Guided Mathematical Reasoning in LLMsYuanhe Zhang, Ilja Kuzborskij, Jason D. Lee, Chenlei Leng et al.ICLR 2026 · 6 citations
- ByteSized32: A Corpus and Challenge Task for Generating Task-Specific World Models Expressed as Text GamesRuoyao Wang, Graham Todd, Xingdi Yuan, Ziang Xiao et al.EMNLP 2023 · 1 citation
- WinoLogic: A Zero-Shot Logic-based Diagnostic Dataset for Winograd Schema ChallengeWeinan He, Canming Huang, Yongmei Liu, Xiaodan ZhuEMNLP 2021 · 7 citations
- FloorplanQA: A Benchmark for Spatial Reasoning in LLMs using Structured RepresentationsFedor Rodionov, Abdelrahman Eldesokey, Michael Birsak, John Femiani et al.ICML 2026 · 14 citations
- Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-ThoughtAbulhair Saparov, He HeICLR 2023 · 38 citations
