Avoiding Side Effects in Complex Environments
Alexander Matt Turner, Neale Ratzlaff, Prasad Tadepalli
Abstract
Reward function specification can be difficult. Rewarding the agent for making a widget may be easy, but penalizing the multitude of possible negative side effects is hard. In toy environments, Attainable Utility Preservation (AUP) avoided side effects by penalizing shifts in the ability to achieve randomly generated goals [22] . We scale this approach to large, randomly generated environments based on Conway's Game of Life. By preserving optimal value for a single randomly generated reward function, AUP incurs modest overhead while leading the agent to complete the specified task and avoid many side effects. Videos and code are available at https://avoiding-side-effects.github.io/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 380 citations
- Exploratory Machine Learning with Unknown UnknownsPeng Zhao, Yu-Jie Zhang, Zhi-Hua ZhouAAAI 2021 · 29 citations
- Quantifying the Sensitivity of Inverse Reinforcement Learning to MisspecificationJoar Max Viktor Skalse, Alessandro AbateICLR 2024 · 5 citations
- Calibrating Conservatism for Scalable OversightWilliam Overman, Mohsen BayatiICML 2026
- O-MAPL: Offline Multi-agent Preference LearningThe Viet Bui, Tien Mai, Thanh Hong NguyenICML 2025
Related papers
- Avoiding Side Effects By Considering Future TasksVictoria Krakovna, Laurent Orseau, Richard Ngo, Miljan Martic et al.NeurIPS 2020 · 54 citations
- Reward Design with Language ModelsMinae Kwon, Sang Michael Xie, Kalesha Bullard, Dorsa SadighICLR 2023 · 21 citations
- Direct Behavior Specification via Constrained Reinforcement LearningJulien Roy, Roger Girgis, Joshua Romoff, Pierre-Luc Bacon et al.ICML 2022 · 46 citations
- Corrigibility Transformation: Constructing Goals That Accept UpdatesRubi HudsonICML 2026
- Behavior Alignment via Reward Function OptimizationDhawal Gupta, Yash Chandak, Scott M. Jordan, Philip S. Thomas et al.NeurIPS 2023 · 27 citations
