Path-Specific Objectives for Safer Agent Incentives
Sebastian Farquhar, Ryan Carey, Tom Everitt
Abstract
We present a general framework for training safe agents whose naive incentives are unsafe. As an example, manipulative or deceptive behaviour can improve rewards but should be avoided. Most approaches fail here: agents maximize expected return by any means necessary. We formally describe settings with `delicate' parts of the state which should not be used as a means to an end. We then train agents to maximize the causal effect of actions on the expected return which is not mediated by the delicate parts of state, using Causal Influence Diagram analysis. The resulting agents have no incentive to control the delicate state. We further show how our framework unifies and generalizes existing proposals.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d8560d6a-e9e9-44bf-a964-e9122993cbefCited by top-tier papers11
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 296 citations
- Honesty Is the Best Policy: Defining and Mitigating AI DeceptionFrancis Ward, Francesca Toni, Francesco Belardinelli, Tom EverittNeurIPS 2023 · 60 citations
- Estimating and Penalizing Induced Preference Shifts in Recommender SystemsMicah D. Carroll, Anca D. Dragan, Stuart Russell, Dylan Hadfield-MenellICML 2022 · 49 citations
- AI Alignment with Changing and Influenceable Reward FunctionsMicah Carroll, Davis Foote, Anand Siththaranjan, Stuart Russell et al.ICML 2024 · 44 citations
- A Complete Criterion for Value of Information in Soluble Influence DiagramsChris van Merwijk, Ryan Carey, Tom EverittAAAI 2022 · 7 citations
Builds on2
- Avoiding Side Effects By Considering Future TasksVictoria Krakovna, Laurent Orseau, Richard Ngo, Miljan Martic et al.NeurIPS 2020 · 54 citations
- Estimating and Penalizing Induced Preference Shifts in Recommender SystemsMicah D. Carroll, Anca D. Dragan, Stuart Russell, Dylan Hadfield-MenellICML 2022 · 49 citations
Related papers
- Agent Incentives: A Causal PerspectiveTom Everitt, Ryan Carey, Eric D. Langlois, Pedro A. Ortega et al.AAAI 2021 · 66 citations
- On the Robustness of Safe Reinforcement Learning under Observational PerturbationsZuxin Liu, Zijian Guo, Zhepeng Cen, Huan Zhang et al.ICLR 2023 · 9 citations
- Reinforcement Learning of Causal Variables Using Mediation AnalysisTue Herlau, Rasmus LarsenAAAI 2022 · 8 citations
- Shield Decentralization for Safe Multi-Agent Reinforcement LearningDaniel Melcer, Christopher Amato, Stavros TripakisNeurIPS 2022 · 26 citations
- Output Supervision Can Obfuscate the Chain of ThoughtJacob Drori, Luke Marks, Bryce Woodworth, Alex Cloud et al.ICLR 2026 · 10 citations
