Lune

ICML2025顶会

The Limits of Predicting Agents from Behaviour

Alexis Bellot, Jonathan Richens, Tom Everitt

出版方
2025年份
2顶会引用

摘要

As the complexity of AI systems and their interactions with the world increases, generating explanations for their behaviour is important for safely deploying AI. For agents, the most natural abstractions for predicting behaviour attribute beliefs, intentions and goals to the system. If an agent behaves as if it has a certain goal or belief, then we can make reasonable predictions about how it will behave in novel situations, including those where comprehensive safety evaluations are untenable. How well can we infer an agent's beliefs from their behaviour, and how reliably can these inferred beliefs predict the agent's behaviour in novel situations? We provide a precise answer to this question under the assumption that the agent's behaviour is guided by a world model. Our contribution is the derivation of novel bounds on the agent's behaviour in new (unseen) deployment environments, which represent a theoretical limit for predicting intentional agents from behavioural data alone. We discuss the implications of these results for several research areas including fairness and safety. 1 Recent research suggests that an AI's behaviour, to the extent that it is consistent with rationality axioms, can be formally described by a (causal) world model (Halpern and Piermont, 2024) . The same conclusion can also be obtained for AIs capable of solving tasks in multiple environments (Richens and Everitt, 2024) . For large language models, there is increasing empirical evidence for the "world model" hypothesis, see e.g.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper2

问问它们各自怎么用它

它引用的顶会 Paper20

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖