Inverse Reinforcement Learning From Like-Minded Teachers
Ritesh Noothigattu, Tom Yan, Ariel D. Procaccia
摘要
We study the problem of learning a policy in a Markov decision process (MDP) based on observations of the actions taken by multiple teachers. We assume that the teachers are like-minded in that their reward functions -- while different from each other -- are random perturbations of an underlying reward function. Under this assumption, we demonstrate that inverse reinforcement learning algorithms that satisfy a certain property -- that of matching feature expectations -- yield policies that are approximately optimal with respect to the underlying reward function, and that no algorithm can do better in the worst case. We also show how to efficiently recover the optimal policy when the MDP has one state -- a setting that is akin to multi-armed bandits.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Policy AggregationParand A. Alamdari, Soroush Ebadian, Ariel D. ProcacciaNeurIPS 2024 · 被引用 11 次
- Envy-free Policy Teaching to Multiple AgentsJiarui Gan, Rupak Majumdar, Adish Singla, Goran RadanovicNeurIPS 2022
相关 Paper
- Inverse Reinforcement Learning from a Gradient-based LearnerGiorgia Ramponi, Gianluca Drappo, Marcello RestelliNeurIPS 2020 · 被引用 16 次
- Identifiability and generalizability from multiple experts in Inverse Reinforcement LearningPaul Rolland, Luca Viano, Norman Schürhoff, Boris Nikolov 等NeurIPS 2022 · 被引用 22 次
- Inverse Reinforcement Learning with the Average Reward CriterionFeiyang Wu, Jingyang Ke, Anqi WuNeurIPS 2023 · 被引用 16 次
- Identifiability in inverse reinforcement learningHaoyang Cao, Samuel N. Cohen, Lukasz SzpruchNeurIPS 2021 · 被引用 72 次
- Sub-optimal Experts mitigate Ambiguity in Inverse Reinforcement LearningRiccardo Poiani, Gabriele Curti, Alberto Maria Metelli, Marcello RestelliNeurIPS 2024 · 被引用 2 次
