The Limits of Predicting Agents from Behaviour
Alexis Bellot, Jonathan Richens, Tom Everitt
Abstract
As the complexity of AI systems and their interactions with the world increases, generating explanations for their behaviour is important for safely deploying AI. For agents, the most natural abstractions for predicting behaviour attribute beliefs, intentions and goals to the system. If an agent behaves as if it has a certain goal or belief, then we can make reasonable predictions about how it will behave in novel situations, including those where comprehensive safety evaluations are untenable. How well can we infer an agent's beliefs from their behaviour, and how reliably can these inferred beliefs predict the agent's behaviour in novel situations? We provide a precise answer to this question under the assumption that the agent's behaviour is guided by a world model. Our contribution is the derivation of novel bounds on the agent's behaviour in new (unseen) deployment environments, which represent a theoretical limit for predicting intentional agents from behavioural data alone. We discuss the implications of these results for several research areas including fairness and safety. 1 Recent research suggests that an AI's behaviour, to the extent that it is consistent with rationality axioms, can be formally described by a (causal) world model (Halpern and Piermont, 2024) . The same conclusion can also be obtained for AIs capable of solving tasks in multiple environments (Richens and Everitt, 2024) . For large language models, there is increasing empirical evidence for the "world model" hypothesis, see e.g.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5f415c94-328d-45bf-9815-816cd00f1c83Cited by top-tier papers2
- A Behavioural and Representational Evaluation of Goal-Directedness in Language Model AgentsRaghu Arghal, Fade Chen, Niall Dalton, Evgenii Kortukov et al.ICML 2026
- World Models in Pieces: Structural Certification for General AgentsYikai Lu, Yifei Wu, Xinyu Lu, Tongxin LiICML 2026
Builds on20
- Causal Abstractions of Neural NetworksAtticus Geiger, Hanson Lu, Thomas Icard, Christopher PottsNeurIPS 2021 · 516 citations
- Language Models Represent Space and TimeWes Gurnee, Max TegmarkICLR 2024 · 303 citations
- Evaluating the World Model Implicit in a Generative ModelKeyon Vafa, Justin Y. Chen, Ashesh Rambachan, Jon M. Kleinberg et al.NeurIPS 2024 · 166 citations
- A Calculus for Stochastic Interventions: Causal Effect Identification and Surrogate ExperimentsJuan D. Correa, Elias BareinboimAAAI 2020 · 90 citations
- Robust agents learn causal world modelsJonathan Richens, Tom EverittICLR 2024 · 78 citations
Related papers
- General agents need world modelsJonathan Richens, Tom Everitt, David AbelICML 2025
- Explainable Reinforcement Learning through a Causal LensPrashan Madumal, Tim Miller, Liz Sonenberg, Frank VetereAAAI 2020 · 408 citations
- Context and Diversity Matter: The Emergence of In-Context Learning in World ModelsFan Wang, ZHIYUAN CHEN, YUXUAN ZHONG, Sunjian Zheng et al.ICLR 2026 · 5 citations
- Generating High-Quality Explanations for Navigation in Partially-Revealed EnvironmentsGregory J. SteinNeurIPS 2021 · 19 citations
- (Mis)Communicating with our AI SystemsLaura Cros Vila, Bob L. T. SturmCHI 2025 · 2 citations
