Policy Gradient Bayesian Robust Optimization for Imitation Learning
Zaynah Javed, Daniel S. Brown, Satvik Sharma, Jerry Zhu, Ashwin Balakrishna, Marek Petrik, Anca D. Dragan, Ken Goldberg
Abstract
The difficulty in specifying rewards for many real-world problems has led to an increased focus on learning rewards from human feedback, such as demonstrations. However, there are often many different reward functions that explain the human feedback, leaving agents with uncertainty over what the true reward function is. While most policy optimization approaches handle this uncertainty by optimizing for expected performance, many applications demand risk-averse behavior. We derive a novel policy gradient-style robust optimization approach, PG-BROIL, that optimizes a soft-robust objective that balances expected performance and risk. To the best of our knowledge, PG-BROIL is the first policy optimization algorithm robust to a distribution of reward hypotheses which can scale to continuous MDPs. Results suggest that PG-BROIL can produce a family of behaviors ranging from risk-neutral to risk-averse and outperforms state-of-the-art imitation learning algorithms when learning from ambiguous demonstrations by hedging against uncertainty, rather than seeking to uniquely identify the demonstrator's reward function.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- The Effects of Reward Misspecification: Mapping and Mitigating Misaligned ModelsAlexander Pan, Kush Bhatia, Jacob SteinhardtICLR 2022 · 293 citations
- Improving Generalization of Alignment with Human Preferences through Group Invariant LearningRui Zheng, Wei Shen, Yuan Hua, Wenbin Lai et al.ICLR 2024 · 25 citations
- Contextual Reliability: When Different Features Matter in Different ContextsGaurav Rohit Ghosal, Amrith Setlur, Daniel S. Brown, Anca D. Dragan et al.ICML 2023 · 3 citations
- Causal Confusion and Reward Misidentification in Preference-Based Reward LearningJeremy Tien, Jerry Zhi-Yang He, Zackory Erickson, Anca D. Dragan et al.ICLR 2023 · 3 citations
- Imitation Beyond Expectation Using Pluralistic Stochastic DominanceAli Farajzadeh, Danyal Saeed, Syed M. Abbas, Rushit N. Shah et al.NeurIPS 2025 · 2 citations
Builds on4
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- Safe Imitation Learning via Fast Bayesian Reward Inference from PreferencesDaniel S. Brown, Russell Coleman, Ravi Srinivasan, Scott NiekumICML 2020 · 113 citations
- Mean-Variance Policy Iteration for Risk-Averse Reinforcement LearningShangtong Zhang, Bo Liu, Shimon WhitesonAAAI 2021 · 44 citations
- Bayesian Robust Optimization for Imitation LearningDaniel S. Brown, Scott Niekum, Marek PetrikNeurIPS 2020 · 43 citations
Related papers
- Distributionally Robust Imitation LearningMohammad Ali Bashiri, Brian D. Ziebart, Xinhua ZhangNeurIPS 2021 · 13 citations
- Learning Utilities from Demonstrations in Markov Decision ProcessesFilippo Lazzati, Alberto Maria MetelliICML 2025
- MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational InferenceRaphaël Baur, Yannick Metz, Maria Gkoulta, Mennatallah El-Assady et al.ICML 2026
- Bounded Risk-Sensitive Markov Games: Forward Policy Design and Inverse Reward Learning with Iterative Reasoning and Cumulative Prospect TheoryRan Tian, Liting Sun, Masayoshi TomizukaAAAI 2021 · 13 citations
- Efficient Exploration of Reward Functions in Inverse Reinforcement Learning via Bayesian OptimizationSreejith Balakrishnan, Quoc Phong Nguyen, Bryan Kian Hsiang Low, Harold SohNeurIPS 2020 · 33 citations
