Parametrically Retargetable Decision-Makers Tend To Seek Power
Alexander Matt Turner, Prasad Tadepalli
Abstract
If capable AI agents are generally incentivized to seek power in service of the objectives we specify for them, then these systems will pose enormous risks, in addition to enormous benefits. In fully observable environments, most reward functions have an optimal policy which seeks power by keeping options open and staying alive [Turner et al., 2021] . However, the real world is neither fully observable, nor must trained agents be even approximately reward-optimal. We consider a range of models of AI decision-making, from optimal, to random, to choices informed by learning and interacting with an environment. We discover that many decision-making functions are retargetable, and that retargetability is sufficient to cause power-seeking tendencies. Our functional criterion is simple and broad. We show that a range of qualitatively dissimilar decision-making procedures incentivize agents to seek power. We demonstrate the flexibility of our results by reasoning about learned policy incentives in Montezuma's Revenge. These results suggest a safety risk: Eventually, retargetable training procedures may train real-world agents which seek power over humans.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 296 citations
- Estimating the Empowerment of Language Model AgentsJinyeop Song, Jeff Gore, Max Kleiman-WeinerICML 2026
Builds on1
Related papers
- Emergent Risk Awareness in Rational Agents under Resource ConstraintsDaniel Jarne Ornia, Nicholas Bishop, Joel Dyer, Wei-Chen Lee et al.NeurIPS 2025 · 5 citations
- Corrigibility Transformation: Constructing Goals That Accept UpdatesRubi HudsonICML 2026
- Utility Theory for Sequential Decision MakingMehran Shakerinava, Siamak RavanbakhshICML 2022 · 8 citations
- Alignment Risks from Capability-Seeking RL TrainingYujun Zhou, Yue Huang, Han Bao, kehan guo et al.ICML 2026
- Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the Machiavelli BenchmarkAlexander Pan, Jun Shern Chan, Andy Zou, Nathaniel Li et al.ICML 2023 · 200 citations
