Avoiding Side Effects By Considering Future Tasks
Victoria Krakovna, Laurent Orseau, Richard Ngo, Miljan Martic, Shane Legg
摘要
Designing reward functions is difficult: the designer has to specify what to do (what it means to complete the task) as well as what not to do (side effects that should be avoided while completing the task). To alleviate the burden on the reward designer, we propose an algorithm to automatically generate an auxiliary reward function that penalizes side effects. This auxiliary objective rewards the ability to complete possible future tasks, which decreases if the agent causes side effects during the current task. The future task reward can also give the agent an incentive to interfere with events in the environment that make future tasks less achievable, such as irreversible actions by other agents. To avoid this interference incentive, we introduce a baseline policy that represents a default course of action (such as doing nothing), and use it to filter out future tasks that are not achievable by default. We formally define interference incentives and show that the future task approach with a baseline policy avoids these incentives in the deterministic case. Using gridworld environments that test for side effects and interference, we show that our method avoids interference and is more effective for avoiding side effects than the common approach of penalizing irreversible actions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- The Limits of Inference Scaling Through ResamplingBenedikt Stroebl, Sayash Kapoor, Arvind NarayananICLR 2026 · 被引用 38 次
- Path-Specific Objectives for Safer Agent IncentivesSebastian Farquhar, Ryan Carey, Tom EverittAAAI 2022 · 被引用 30 次
- Quantifying the Sensitivity of Inverse Reinforcement Learning to MisspecificationJoar Max Viktor Skalse, Alessandro AbateICLR 2024 · 被引用 5 次
- Calibrating Conservatism for Scalable OversightWilliam Overman, Mohsen BayatiICML 2026
相关 Paper
- Behavior Alignment via Reward Function OptimizationDhawal Gupta, Yash Chandak, Scott M. Jordan, Philip S. Thomas 等NeurIPS 2023 · 被引用 27 次
- Avoiding Side Effects in Complex EnvironmentsAlexander Matt Turner, Neale Ratzlaff, Prasad TadepalliNeurIPS 2020 · 被引用 40 次
- Admissible Policy Teaching through Reward DesignKiarash Banihashem, Adish Singla, Jiarui Gan, Goran RadanovicAAAI 2022 · 被引用 18 次
- Querying to Find a Safe Policy under Uncertain Safety Constraints in Markov Decision ProcessesShun Zhang, Edmund H. Durfee, Satinder SinghAAAI 2020 · 被引用 6 次
- On the Expressivity of Markov RewardDavid Abel, Will Dabney, Anna Harutyunyan, Mark K. Ho 等NeurIPS 2021 · 被引用 107 次
