Safe Imitation Learning via Fast Bayesian Reward Inference from Preferences
Daniel S. Brown, Russell Coleman, Ravi Srinivasan, Scott Niekum
摘要
Bayesian reward learning from demonstrations enables rigorous safety and uncertainty analysis when performing imitation learning. However, Bayesian reward learning methods are typically computationally intractable for complex control problems. We propose Bayesian Reward Extrapolation (Bayesian REX), a highly efficient Bayesian reward learning algorithm that scales to high-dimensional imitation learning problems by pre-training a low-dimensional feature encoding via self-supervised tasks and then leveraging preferences over demonstrations to perform fast Bayesian inference. Bayesian REX can learn to play Atari games from demonstrations, without access to the game score and can generate 100,000 samples from the posterior over reward functions in only 5 minutes on a personal laptop. Bayesian REX also results in imitation learning performance that is competitive with or better than state-of-the-art methods that only learn point estimates of the reward function. Finally, Bayesian REX enables efficient high-confidence policy evaluation without having access to samples of the reward function. These high-confidence performance bounds can be used to rank the performance and risk of a variety of evaluation policies and provide a way to detect reward hacking behaviors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- The Effects of Reward Misspecification: Mapping and Mitigating Misaligned ModelsAlexander Pan, Kush Bhatia, Jacob SteinhardtICLR 2022 · 被引用 293 次
- Inverse Preference Learning: Preference-based RL without a Reward FunctionJoey Hejna, Dorsa SadighNeurIPS 2023 · 被引用 92 次
- Bayesian Robust Optimization for Imitation LearningDaniel S. Brown, Scott Niekum, Marek PetrikNeurIPS 2020 · 被引用 43 次
- The MAGICAL Benchmark for Robust ImitationSam Toyer, Rohin Shah, Andrew Critch, Stuart RussellNeurIPS 2020 · 被引用 42 次
- Value Alignment VerificationDaniel S. Brown, Jordan Schneider, Anca D. Dragan, Scott NiekumICML 2021 · 被引用 41 次
相关 Paper
- Scalable Bayesian Inverse Reinforcement LearningAlex James Chan, Mihaela van der SchaarICLR 2021 · 被引用 11 次
- Intrinsic Reward Driven Imitation Learning via Generative ModelXingrui Yu, Yueming Lyu, Ivor W. TsangICML 2020 · 被引用 63 次
- A Hierarchical Bayesian Approach to Inverse Reinforcement Learning with Symbolic Reward MachinesWeichao Zhou, Wenchao LiICML 2022 · 被引用 15 次
- Visual Imitation Learning with Patch RewardsMinghuan Liu, Tairan He, Weinan Zhang, Shuicheng Yan 等ICLR 2023 · 被引用 1 次
- Causal Confusion and Reward Misidentification in Preference-Based Reward LearningJeremy Tien, Jerry Zhi-Yang He, Zackory Erickson, Anca D. Dragan 等ICLR 2023 · 被引用 3 次
