Path-Specific Objectives for Safer Agent Incentives
Sebastian Farquhar, Ryan Carey, Tom Everitt
摘要
We present a general framework for training safe agents whose naive incentives are unsafe. As an example, manipulative or deceptive behaviour can improve rewards but should be avoided. Most approaches fail here: agents maximize expected return by any means necessary. We formally describe settings with `delicate' parts of the state which should not be used as a means to an end. We then train agents to maximize the causal effect of actions on the expected return which is not mediated by the delicate parts of state, using Causal Influence Diagram analysis. The resulting agents have no incentive to control the delicate state. We further show how our framework unifies and generalizes existing proposals.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 被引用 296 次
- Honesty Is the Best Policy: Defining and Mitigating AI DeceptionFrancis Ward, Francesca Toni, Francesco Belardinelli, Tom EverittNeurIPS 2023 · 被引用 60 次
- Estimating and Penalizing Induced Preference Shifts in Recommender SystemsMicah D. Carroll, Anca D. Dragan, Stuart Russell, Dylan Hadfield-MenellICML 2022 · 被引用 49 次
- AI Alignment with Changing and Influenceable Reward FunctionsMicah Carroll, Davis Foote, Anand Siththaranjan, Stuart Russell 等ICML 2024 · 被引用 44 次
- A Complete Criterion for Value of Information in Soluble Influence DiagramsChris van Merwijk, Ryan Carey, Tom EverittAAAI 2022 · 被引用 7 次
它引用的顶会 Paper2
- Avoiding Side Effects By Considering Future TasksVictoria Krakovna, Laurent Orseau, Richard Ngo, Miljan Martic 等NeurIPS 2020 · 被引用 54 次
- Estimating and Penalizing Induced Preference Shifts in Recommender SystemsMicah D. Carroll, Anca D. Dragan, Stuart Russell, Dylan Hadfield-MenellICML 2022 · 被引用 49 次
相关 Paper
- Agent Incentives: A Causal PerspectiveTom Everitt, Ryan Carey, Eric D. Langlois, Pedro A. Ortega 等AAAI 2021 · 被引用 66 次
- On the Robustness of Safe Reinforcement Learning under Observational PerturbationsZuxin Liu, Zijian Guo, Zhepeng Cen, Huan Zhang 等ICLR 2023 · 被引用 9 次
- Reinforcement Learning of Causal Variables Using Mediation AnalysisTue Herlau, Rasmus LarsenAAAI 2022 · 被引用 8 次
- Shield Decentralization for Safe Multi-Agent Reinforcement LearningDaniel Melcer, Christopher Amato, Stavros TripakisNeurIPS 2022 · 被引用 26 次
- Output Supervision Can Obfuscate the Chain of ThoughtJacob Drori, Luke Marks, Bryce Woodworth, Alex Cloud 等ICLR 2026 · 被引用 10 次
