Safe Imitation Learning via Fast Bayesian Reward Inference from Preferences
Daniel S. Brown, Russell Coleman, Ravi Srinivasan, Scott Niekum
Abstract
Bayesian reward learning from demonstrations enables rigorous safety and uncertainty analysis when performing imitation learning. However, Bayesian reward learning methods are typically computationally intractable for complex control problems. We propose Bayesian Reward Extrapolation (Bayesian REX), a highly efficient Bayesian reward learning algorithm that scales to high-dimensional imitation learning problems by pre-training a low-dimensional feature encoding via self-supervised tasks and then leveraging preferences over demonstrations to perform fast Bayesian inference. Bayesian REX can learn to play Atari games from demonstrations, without access to the game score and can generate 100,000 samples from the posterior over reward functions in only 5 minutes on a personal laptop. Bayesian REX also results in imitation learning performance that is competitive with or better than state-of-the-art methods that only learn point estimates of the reward function. Finally, Bayesian REX enables efficient high-confidence policy evaluation without having access to samples of the reward function. These high-confidence performance bounds can be used to rank the performance and risk of a variety of evaluation policies and provide a way to detect reward hacking behaviors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b1d2ee73-5466-4d54-9606-798650da075dCited by top-tier papers24
- The Effects of Reward Misspecification: Mapping and Mitigating Misaligned ModelsAlexander Pan, Kush Bhatia, Jacob SteinhardtICLR 2022 · 293 citations
- Inverse Preference Learning: Preference-based RL without a Reward FunctionJoey Hejna, Dorsa SadighNeurIPS 2023 · 92 citations
- Bayesian Robust Optimization for Imitation LearningDaniel S. Brown, Scott Niekum, Marek PetrikNeurIPS 2020 · 43 citations
- The MAGICAL Benchmark for Robust ImitationSam Toyer, Rohin Shah, Andrew Critch, Stuart RussellNeurIPS 2020 · 42 citations
- Value Alignment VerificationDaniel S. Brown, Jordan Schneider, Anca D. Dragan, Scott NiekumICML 2021 · 41 citations
Related papers
- Scalable Bayesian Inverse Reinforcement LearningAlex James Chan, Mihaela van der SchaarICLR 2021 · 11 citations
- Intrinsic Reward Driven Imitation Learning via Generative ModelXingrui Yu, Yueming Lyu, Ivor W. TsangICML 2020 · 63 citations
- A Hierarchical Bayesian Approach to Inverse Reinforcement Learning with Symbolic Reward MachinesWeichao Zhou, Wenchao LiICML 2022 · 15 citations
- Visual Imitation Learning with Patch RewardsMinghuan Liu, Tairan He, Weinan Zhang, Shuicheng Yan et al.ICLR 2023 · 1 citation
- Causal Confusion and Reward Misidentification in Preference-Based Reward LearningJeremy Tien, Jerry Zhi-Yang He, Zackory Erickson, Anca D. Dragan et al.ICLR 2023 · 3 citations
