Bayesian Robust Optimization for Imitation Learning
Daniel S. Brown, Scott Niekum, Marek Petrik
摘要
One of the main challenges in imitation learning is determining what action an agent should take when outside the state distribution of the demonstrations. Inverse reinforcement learning (IRL) can enable generalization to new states by learning a parameterized reward function, but these approaches still face uncertainty over the true reward function and corresponding optimal policy. Existing safe imitation learning approaches based on IRL deal with this uncertainty using a maxmin framework that optimizes a policy under the assumption of an adversarial reward function, whereas risk-neutral IRL approaches either optimize a policy for the mean or MAP reward function. While completely ignoring risk can lead to overly aggressive and unsafe policies, optimizing in a fully adversarial sense is also problematic as it can lead to overly conservative policies that perform poorly in practice. To provide a bridge between these two extremes, we propose Bayesian Robust Optimization for Imitation Learning (BROIL). BROIL leverages Bayesian reward function inference and a user specific risk tolerance to efficiently optimize a robust policy that balances expected return and conditional value at risk. Our empirical results show that BROIL provides a natural way to interpolate between return-maximizing and risk-minimizing behaviors and outperforms existing risksensitive and risk-neutral inverse reinforcement learning algorithms. Code is available at https://github.com/dsbrown1331/broil .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Adversarial training for high-stakes reliabilityDaniel M. Ziegler, Seraphina Nix, Lawrence Chan, Tim Bauman 等NeurIPS 2022 · 被引用 79 次
- Universal Off-Policy EvaluationYash Chandak, Scott Niekum, Bruno C. da Silva, Erik G. Learned-Miller 等NeurIPS 2021 · 被引用 64 次
- Multi-Modal Inverse Constrained Reinforcement Learning from a Mixture of DemonstrationsGuanren Qiao, Guiliang Liu, Pascal Poupart, Zhiqiang XuNeurIPS 2023 · 被引用 28 次
- Improving Generalization of Alignment with Human Preferences through Group Invariant LearningRui Zheng, Wei Shen, Yuan Hua, Wenbin Lai 等ICLR 2024 · 被引用 25 次
- Policy Gradient Bayesian Robust Optimization for Imitation LearningZaynah Javed, Daniel S. Brown, Satvik Sharma, Jerry Zhu 等ICML 2021 · 被引用 18 次
它引用的顶会 Paper3
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Decoupling Representation Learning from Reinforcement LearningAdam Stooke, Kimin Lee, Pieter Abbeel, Michael LaskinICML 2021 · 被引用 389 次
- Safe Imitation Learning via Fast Bayesian Reward Inference from PreferencesDaniel S. Brown, Russell Coleman, Ravi Srinivasan, Scott NiekumICML 2020 · 被引用 113 次
相关 Paper
- Distributionally Robust Imitation LearningMohammad Ali Bashiri, Brian D. Ziebart, Xinhua ZhangNeurIPS 2021 · 被引用 13 次
- Efficient Exploration of Reward Functions in Inverse Reinforcement Learning via Bayesian OptimizationSreejith Balakrishnan, Quoc Phong Nguyen, Bryan Kian Hsiang Low, Harold SohNeurIPS 2020 · 被引用 33 次
- BC-IRL: Learning Generalizable Reward Functions from DemonstrationsAndrew Szot, Amy Zhang, Dhruv Batra, Zsolt Kira 等ICLR 2023 · 被引用 1 次
- Learning Utilities from Demonstrations in Markov Decision ProcessesFilippo Lazzati, Alberto Maria MetelliICML 2025
- Trust Region Inverse Reinforcement Learning: Explicit Dual Ascent using Local Policy UpdatesAnish Abhijit Diwan, Davide Tateo, Christopher Mower, Haitham Bou Ammar 等ICML 2026
