Identifiability in inverse reinforcement learning
Haoyang Cao, Samuel N. Cohen, Lukasz Szpruch
摘要
Inverse reinforcement learning attempts to reconstruct the reward function in a Markov decision problem, using observations of agent actions. As already observed in Russell [1998] the problem is ill-posed, and the reward function is not identifiable, even under the presence of perfect information about optimal behavior. We provide a resolution to this non-identifiability for problems with entropy regularization. For a given environment, we fully characterize the reward functions leading to a given policy and demonstrate that, given demonstrations of actions for the same reward under two distinct discount factors, or under sufficiently different environments, the unobserved reward can be recovered up to a constant. We also give general necessary and sufficient conditions for reconstruction of time-homogeneous rewards on finite horizons, and for action-independent rewards, generalizing recent results of Kim et al. [2021] and Fu et al. [2018] .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Maximum-Likelihood Inverse Reinforcement Learning with Finite-Time GuaranteesSiliang Zeng, Chenliang Li, Alfredo García, Mingyi HongNeurIPS 2022 · 被引用 60 次
- Invariance in Policy Optimisation and Partial Identifiability in Reward LearningJoar Max Viktor Skalse, Matthew Farrugia-Roberts, Stuart Russell, Alessandro Abate 等ICML 2023 · 被引用 56 次
- When Demonstrations meet Generative World Models: A Maximum Likelihood Framework for Offline Inverse Reinforcement LearningSiliang Zeng, Chenliang Li, Alfredo García, Mingyi HongNeurIPS 2023 · 被引用 33 次
- Coherent Soft Imitation LearningJoe Watson, Sandy H. Huang, Nicolas HeessNeurIPS 2023 · 被引用 26 次
- Identifiability and generalizability from multiple experts in Inverse Reinforcement LearningPaul Rolland, Luca Viano, Norman Schürhoff, Boris Nikolov 等NeurIPS 2022 · 被引用 22 次
它引用的顶会 Paper1
相关 Paper
- Decoding Rewards in Competitive Games: Inverse Game Theory with Entropy RegularizationJunyi Liao, Zihan Zhu, Ethan X. Fang, Zhuoran Yang 等ICML 2025
- Identifiability and Generalizability in Constrained Inverse Reinforcement LearningAndreas Schlaginhaufen, Maryam KamgarpourICML 2023 · 被引用 18 次
- On Feasible Rewards in Multi-Agent Inverse Reinforcement LearningTill Freihaut, Giorgia RamponiNeurIPS 2025 · 被引用 5 次
- Maximum Likelihood Constraint Inference for Inverse Reinforcement LearningDexter R. R. Scobee, S. Shankar SastryICLR 2020 · 被引用 74 次
- Inverse Reinforcement Learning from a Gradient-based LearnerGiorgia Ramponi, Gianluca Drappo, Marcello RestelliNeurIPS 2020 · 被引用 16 次
