Behavior Alignment via Reward Function Optimization
Dhawal Gupta, Yash Chandak, Scott M. Jordan, Philip S. Thomas, Bruno C. da Silva
摘要
Designing reward functions for efficiently guiding reinforcement learning (RL) agents toward specific behaviors is a complex task. This is challenging since it requires the identification of reward structures that are not sparse and that avoid inadvertently inducing undesirable behaviors. Naively modifying the reward structure to offer denser and more frequent feedback can lead to unintended outcomes and promote behaviors that are not aligned with the designer's intended goal. Although potential-based reward shaping is often suggested as a remedy, we systematically investigate settings where deploying it often significantly impairs performance. To address these issues, we introduce a new framework that uses a bi-level objective to learn behavior alignment reward functions. These functions integrate auxiliary rewards reflecting a designer's heuristics and domain knowledge with the environment's primary rewards. Our approach automatically determines the most effective way to blend these types of feedback, thereby enhancing robustness against heuristic reward misspecification. Remarkably, it can also adapt an agent's policy optimization process to mitigate suboptimalities resulting from limitations and biases inherent in the underlying RL algorithms. We evaluate our method's efficacy on a diverse set of tasks, from small-scale experiments to high-dimensional control challenges. We investigate heuristic auxiliary rewards of varying quality -- some of which are beneficial and others detrimental to the learning process. Our results show that our framework offers a robust and principled way to integrate designer-specified heuristics. It not only addresses key shortcomings of existing approaches but also consistently leads to high-performing solutions, even when given misaligned or poorly-specified auxiliary reward functions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Reward Shaping for Reinforcement Learning with An Assistant Reward AgentHaozhe Ma, Kuankuan Sima, Thanh Vinh Vo, Di Fu 等ICML 2024 · 被引用 34 次
- Solving Minimum-Cost Reach Avoid using Reinforcement LearningOswin So, Cheng Ge, Chuchu FanNeurIPS 2024 · 被引用 24 次
- MRO: Enhancing Reasoning in Diffusion Language Models via Multi-Reward OptimizationChenglong Wang, Yang Gan, Hang Zhou, Chi Hu 等NeurIPS 2025 · 被引用 4 次
- Learning to Stabilize Online Reinforcement Learning in Unbounded State SpacesBrahma S. Pavse, Matthew Zurek, Yudong Chen, Qiaomin Xie 等ICML 2024 · 被引用 3 次
- Going Beyond Heuristics by Imposing Policy Improvement as a ConstraintChi-Chang Lee, Zhang-Wei Hong, Pulkit AgrawalNeurIPS 2024 · 被引用 2 次
它引用的顶会 Paper8
- Differentiation of Blackbox Combinatorial SolversMarin Vlastelica Pogancic, Anselm Paulus, Vít Musil, Georg Martius 等ICLR 2020 · 被引用 341 次
- Learning to Utilize Shaping Rewards: A New Approach of Reward ShapingYujing Hu, Weixun Wang, Hangtian Jia, Yixiang Wang 等NeurIPS 2020 · 被引用 256 次
- Discovering Reinforcement Learning AlgorithmsJunhyuk Oh, Matteo Hessel, Wojciech M. Czarnecki, Zhongwen Xu 等NeurIPS 2020 · 被引用 154 次
- A Game Theoretic Framework for Model Based Reinforcement LearningAravind Rajeswaran, Igor Mordatch, Vikash KumarICML 2020 · 被引用 137 次
- Improving Generalization in Meta Reinforcement Learning using Learned ObjectivesLouis Kirsch, Sjoerd van Steenkiste, Jürgen SchmidhuberICLR 2020 · 被引用 132 次
相关 Paper
- Centralized Reward Agent for Knowledge Sharing and Transfer in Multi-Task Reinforcement LearningHaozhe Ma, Zhengding Luo, Thanh Vinh Vo, Kuankuan Sima 等NeurIPS 2025 · 被引用 9 次
- Programmatic Reward Design by ExampleWeichao Zhou, Wenchao LiAAAI 2022 · 被引用 15 次
- Automatic Reward Shaping from Confounded Offline DataMingxuan Li, Junzhe Zhang, Elias BareinboimICML 2025
- BAMDP Shaping: a Unified Framework for Intrinsic Motivation and Reward ShapingAly Lidayan, Michael D. Dennis, Stuart RussellICLR 2025
- The Perils of Trial-and-Error Reward Design: Misdesign through Overfitting and Invalid Task SpecificationsSerena Booth, W. Bradley Knox, Julie Shah, Scott Niekum 等AAAI 2023 · 被引用 103 次
