A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment
Raanan Yehezkel Rohekar, Yaniv Gurwicz, Sungduk Yu, Estelle Aflalo, Vasudev Lal
Abstract
Are generative pre-trained transformer (GPT) models, trained only to predict the next token, implicitly learning a world model from which sequences are generated one token at a time? We address this question by deriving a causal interpretation of the attention mechanism in GPT and presenting a causal world model that arises from this interpretation. Furthermore, we propose that GPT models, at inference time, can be utilized for zeroshot causal structure learning for input sequences, and introduce a corresponding confidence score. Empirical tests were conducted in controlled environments using the setups of the Othello and Chess strategy games. A GPT, pre-trained on realworld games played with the intention of winning, was tested on out-of-distribution synthetic data consisting of sequences of random legal moves. We find that the GPT model is likely to generate legal next moves for out-of-distribution sequences for which a causal structure is encoded in the attention mechanism with high confidence. In cases where it generates illegal moves, it also fails to capture a causal structure.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5bdfc2ec-e195-456a-a609-265bc4cd5a92Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Are Emergent Abilities of Large Language Models a Mirage?Rylan Schaeffer, Brando Miranda, Sanmi KoyejoNeurIPS 2023 · 796 citations
- Chess as a Testbed for Language Model State TrackingShubham Toshniwal, Sam Wiseman, Karen Livescu, Kevin GimpelAAAI 2022 · 77 citations
- Causal Interpretation of Self-Attention in Pre-Trained TransformersRaanan Y. Rohekar, Yaniv Gurwicz, Shami NisimovNeurIPS 2023 · 62 citations
- Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic TaskKenneth Li, Aspen K. Hopkins, David Bau, Fernanda B. Viégas et al.ICLR 2023 · 60 citations
Related papers
- Causal Discovery and Inference through Next-Token PredictionEivinas Butkus, Nikolaus KriegeskorteNeurIPS 2025 · 3 citations
- A Fixed-Point Approach for Causal Generative ModelingMeyer Scetbon, Joel Jennings, Agrin Hilmkil, Cheng Zhang et al.ICML 2024 · 4 citations
- How Transformers Learn Causal Structure with Gradient DescentEshaan Nichani, Alex Damian, Jason D. LeeICML 2024 · 117 citations
- Verification of the Implicit World Model in a Generative Model via Adversarial SequencesAndrás Balogh, Márk JelasityICLR 2026 · 1 citation
- Humanoid Locomotion as Next Token PredictionIlija Radosavovic, Bike Zhang, Baifeng Shi, Jathushan Rajasegaran et al.NeurIPS 2024 · 128 citations
