Inverse Reinforcement Learning From Like-Minded Teachers
Ritesh Noothigattu, Tom Yan, Ariel D. Procaccia
Abstract
We study the problem of learning a policy in a Markov decision process (MDP) based on observations of the actions taken by multiple teachers. We assume that the teachers are like-minded in that their reward functions -- while different from each other -- are random perturbations of an underlying reward function. Under this assumption, we demonstrate that inverse reinforcement learning algorithms that satisfy a certain property -- that of matching feature expectations -- yield policies that are approximately optimal with respect to the underlying reward function, and that no algorithm can do better in the worst case. We also show how to efficiently recover the optimal policy when the MDP has one state -- a setting that is akin to multi-armed bandits.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9deaa079-176f-4f71-b2c4-eed714449f0cCited by top-tier papers2
- Policy AggregationParand A. Alamdari, Soroush Ebadian, Ariel D. ProcacciaNeurIPS 2024 · 11 citations
- Envy-free Policy Teaching to Multiple AgentsJiarui Gan, Rupak Majumdar, Adish Singla, Goran RadanovicNeurIPS 2022
Related papers
- Inverse Reinforcement Learning from a Gradient-based LearnerGiorgia Ramponi, Gianluca Drappo, Marcello RestelliNeurIPS 2020 · 16 citations
- Identifiability and generalizability from multiple experts in Inverse Reinforcement LearningPaul Rolland, Luca Viano, Norman Schürhoff, Boris Nikolov et al.NeurIPS 2022 · 22 citations
- Inverse Reinforcement Learning with the Average Reward CriterionFeiyang Wu, Jingyang Ke, Anqi WuNeurIPS 2023 · 16 citations
- Identifiability in inverse reinforcement learningHaoyang Cao, Samuel N. Cohen, Lukasz SzpruchNeurIPS 2021 · 72 citations
- Sub-optimal Experts mitigate Ambiguity in Inverse Reinforcement LearningRiccardo Poiani, Gabriele Curti, Alberto Maria Metelli, Marcello RestelliNeurIPS 2024 · 2 citations
