Measuring Goal-Directedness
Matt MacDermott, James Fox, Francesco Belardinelli, Tom Everitt
摘要
We define maximum entropy goal-directedness (MEG), a formal measure of goaldirectedness in causal models and Markov decision processes, and give algorithms for computing it. Measuring goal-directedness is important, as it is a critical element of many concerns about harm from AI. It is also of philosophical interest, as goal-directedness is a key aspect of agency. MEG is based on an adaptation of the maximum causal entropy framework used in inverse reinforcement learning. It can measure goal-directedness with respect to a known utility function, a hypothesis class of utility functions, or a set of random variables. We prove that MEG satisfies several desiderata and demonstrate our algorithms with small-scale experiments 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- A Principle of Targeted Intervention for Multi-Agent Reinforcement LearningAnjie Liu, Jianhong Wang, Samuel Kaski, Jun Wang 等NeurIPS 2025 · 被引用 3 次
- A Behavioural and Representational Evaluation of Goal-Directedness in Language Model AgentsRaghu Arghal, Fade Chen, Niall Dalton, Evgenii Kortukov 等ICML 2026
它引用的顶会 Paper5
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 被引用 296 次
- Goal Misgeneralization in Deep Reinforcement LearningLauro Langosco di Langosco, Jack Koch, Lee D. Sharkey, Jacob Pfau 等ICML 2022 · 被引用 128 次
- Robust agents learn causal world modelsJonathan Richens, Tom EverittICLR 2024 · 被引用 78 次
- Agent Incentives: A Causal PerspectiveTom Everitt, Ryan Carey, Eric D. Langlois, Pedro A. Ortega 等AAAI 2021 · 被引用 66 次
- Honesty Is the Best Policy: Defining and Mitigating AI DeceptionFrancis Ward, Francesca Toni, Francesco Belardinelli, Tom EverittNeurIPS 2023 · 被引用 60 次
相关 Paper
- Plasticity as the Mirror of EmpowermentDavid Abel, Michael Bowling, André Barreto, Will Dabney 等NeurIPS 2025 · 被引用 9 次
- Reward Identification in Inverse Reinforcement LearningKuno Kim, Shivam Garg, Kirankumar Shiragur, Stefano ErmonICML 2021 · 被引用 43 次
- Maximum Causal Entropy Specification Inference from DemonstrationsMarcell Vazquez-Chanlatte, Sanjit A. SeshiaCAV 2020 · 被引用 8 次
- Maximum Likelihood Constraint Inference for Inverse Reinforcement LearningDexter R. R. Scobee, S. Shankar SastryICLR 2020 · 被引用 74 次
- On Measuring Influence in Avoiding Undesired FutureLue Tao, Tian-Zuo Wang, Yuan Jiang, Zhi-Hua ZhouICLR 2026
