Parametrically Retargetable Decision-Makers Tend To Seek Power
Alexander Matt Turner, Prasad Tadepalli
摘要
If capable AI agents are generally incentivized to seek power in service of the objectives we specify for them, then these systems will pose enormous risks, in addition to enormous benefits. In fully observable environments, most reward functions have an optimal policy which seeks power by keeping options open and staying alive [Turner et al., 2021] . However, the real world is neither fully observable, nor must trained agents be even approximately reward-optimal. We consider a range of models of AI decision-making, from optimal, to random, to choices informed by learning and interacting with an environment. We discover that many decision-making functions are retargetable, and that retargetability is sufficient to cause power-seeking tendencies. Our functional criterion is simple and broad. We show that a range of qualitatively dissimilar decision-making procedures incentivize agents to seek power. We demonstrate the flexibility of our results by reasoning about learned policy incentives in Montezuma's Revenge. These results suggest a safety risk: Eventually, retargetable training procedures may train real-world agents which seek power over humans.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 被引用 296 次
- Estimating the Empowerment of Language Model AgentsJinyeop Song, Jeff Gore, Max Kleiman-WeinerICML 2026
它引用的顶会 Paper1
相关 Paper
- Emergent Risk Awareness in Rational Agents under Resource ConstraintsDaniel Jarne Ornia, Nicholas Bishop, Joel Dyer, Wei-Chen Lee 等NeurIPS 2025 · 被引用 5 次
- Corrigibility Transformation: Constructing Goals That Accept UpdatesRubi HudsonICML 2026
- Utility Theory for Sequential Decision MakingMehran Shakerinava, Siamak RavanbakhshICML 2022 · 被引用 8 次
- Alignment Risks from Capability-Seeking RL TrainingYujun Zhou, Yue Huang, Han Bao, kehan guo 等ICML 2026
- Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the Machiavelli BenchmarkAlexander Pan, Jun Shern Chan, Andy Zou, Nathaniel Li 等ICML 2023 · 被引用 200 次
