Avoiding Side Effects in Complex Environments
Alexander Matt Turner, Neale Ratzlaff, Prasad Tadepalli
摘要
Reward function specification can be difficult. Rewarding the agent for making a widget may be easy, but penalizing the multitude of possible negative side effects is hard. In toy environments, Attainable Utility Preservation (AUP) avoided side effects by penalizing shifts in the ability to achieve randomly generated goals [22] . We scale this approach to large, randomly generated environments based on Conway's Game of Life. By preserving optimal value for a single randomly generated reward function, AUP incurs modest overhead while leading the agent to complete the specified task and avoid many side effects. Videos and code are available at https://avoiding-side-effects.github.io/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 被引用 380 次
- Exploratory Machine Learning with Unknown UnknownsPeng Zhao, Yu-Jie Zhang, Zhi-Hua ZhouAAAI 2021 · 被引用 29 次
- Quantifying the Sensitivity of Inverse Reinforcement Learning to MisspecificationJoar Max Viktor Skalse, Alessandro AbateICLR 2024 · 被引用 5 次
- Calibrating Conservatism for Scalable OversightWilliam Overman, Mohsen BayatiICML 2026
- O-MAPL: Offline Multi-agent Preference LearningThe Viet Bui, Tien Mai, Thanh Hong NguyenICML 2025
相关 Paper
- Avoiding Side Effects By Considering Future TasksVictoria Krakovna, Laurent Orseau, Richard Ngo, Miljan Martic 等NeurIPS 2020 · 被引用 54 次
- Reward Design with Language ModelsMinae Kwon, Sang Michael Xie, Kalesha Bullard, Dorsa SadighICLR 2023 · 被引用 21 次
- Direct Behavior Specification via Constrained Reinforcement LearningJulien Roy, Roger Girgis, Joshua Romoff, Pierre-Luc Bacon 等ICML 2022 · 被引用 46 次
- Corrigibility Transformation: Constructing Goals That Accept UpdatesRubi HudsonICML 2026
- Behavior Alignment via Reward Function OptimizationDhawal Gupta, Yash Chandak, Scott M. Jordan, Philip S. Thomas 等NeurIPS 2023 · 被引用 27 次
