What Has a Foundation Model Found? Using Inductive Bias to Probe for World Models
Keyon Vafa, Peter G. Chang, Ashesh Rambachan, Sendhil Mullainathan
Abstract
Foundation models are premised on the idea that sequence prediction can uncover deeper domain understanding, much like how Kepler's predictions of planetary motion later led to the discovery of Newtonian mechanics. However, evaluating whether these models truly capture deeper structure remains a challenge. We develop a technique for evaluating foundation models that examines how they adapt to synthetic datasets generated from some postulated world model. Our technique measures whether the foundation model's inductive bias aligns with the world model, and so we refer to it as an inductive bias probe. Across multiple domains, we find that foundation models can excel at their training tasks yet fail to develop inductive biases towards the underlying world model when adapted to new tasks. We particularly find that foundation models trained on orbital trajectories consistently fail to apply Newtonian mechanics when adapted to new physics tasks. Further analysis reveals that these models behave as if they develop task-specific heuristics that fail to generalize.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5f08ef29-5c6d-47db-abb6-76d089408decCited by top-tier papers8
- Language Models Struggle to Use Representations Learned In-ContextMichael A. Lepori, Tal Linzen, Ann Yuan, Katja FilippovaACL 2026 · 3 citations
- Softplus Attention with Re-weighting Boosts Length Extrapolation in Large Language ModelsBo Gao, Michael Spratling, Letizia GionfridaICML 2026 · 1 citation
- Learning Task-Sufficient World Models by Synergizing Agentic Exploration and Structured ModelingFan Feng, Yujia Zheng, Minghao Fu, Yongqiang Chen et al.ICML 2026
- Convergent World Representations and Divergent TasksCore Francisco ParkICML 2026
- Learning to Extrapolate to New Tasks: A Relational Approach to Task ExtrapolationAdam Ousherovitch, Yixin WangICML 2026
Builds on17
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 1,407 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Discovering Symbolic Models from Deep Learning with Inductive BiasesMiles D. Cranmer, Alvaro Sanchez-Gonzalez, Peter W. Battaglia, Rui Xu et al.NeurIPS 2020 · 736 citations
- A Meta-Transfer Objective for Learning to Disentangle Causal MechanismsYoshua Bengio, Tristan Deleu, Nasim Rahaman, Nan Rosemary Ke et al.ICLR 2020 · 371 citations
- Language Models Represent Space and TimeWes Gurnee, Max TegmarkICLR 2024 · 303 citations
Related papers
- Zero-shot forecasting of chaotic systemsYuanzhao Zhang, William GilpinICLR 2025
- Understanding the Implicit Biases of Design Choices for Time Series Foundation ModelsAnnan Yu, Danielle C. Maddix, Boran Han, Xiyuan Zhang et al.ICLR 2026 · 11 citations
- From Kepler to Newton: Inductive Biases Guide Learned World Models in TransformersZiming Liu, Surya Ganguli, Andreas ToliasICML 2026
- In-Context Fine-Tuning for Time-Series Foundation ModelsMatthew Faw, Rajat Sen, Yichen Zhou, Abhimanyu DasICML 2025
- The Perception–Physics Paradox: Probing Scientific Alignment with TC-BenchDingling Yao, Andrea Polesello, Adeel Pervez, Caroline Muller et al.ICML 2026
