Quantifying the Sensitivity of Inverse Reinforcement Learning to Misspecification
Joar Max Viktor Skalse, Alessandro Abate
Abstract
Inverse reinforcement learning (IRL) aims to infer an agent's preferences (represented as a reward function ) from their behaviour (represented as a policy ). To do this, we need a behavioural model of how relates to . In the current literature, the most common behavioural models are optimality, Boltzmann-rationality, and causal entropy maximisation. However, the true relationship between a human's preferences and their behaviour is much more complex than any of these behavioural models. This means that the behavioural models are misspecified, which raises the concern that they may lead to systematic errors if applied to real data. In this paper, we analyse how sensitive the IRL problem is to misspecification of the behavioural model. Specifically, we provide necessary and sufficient conditions that completely characterise how the observed data may differ from the assumed behavioural model without incurring an error above a given threshold. In addition to this, we also characterise the conditions under which a behavioural model is robust to small perturbations of the observed policy, and we analyse how robust many behavioural models are to misspecification of their parameter values (such as e.g. the discount rate). Our analysis suggests that the IRL problem is highly sensitive to misspecification, in the sense that very mild misspecification can lead to very large errors in the inferred reward function.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 95a1e9c0-7d83-43bc-8950-c984565ea227Cited by top-tier papers5
- Modeling Others' Minds as CodeKunal Jha, Aydan Yuenan Huang, Eric Ye, Natasha Jaques et al.ICLR 2026 · 6 citations
- Learning Utilities from Demonstrations in Markov Decision ProcessesFilippo Lazzati, Alberto Maria MetelliICML 2025
- The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low RegretLukas Fluri, Leon Lang, Alessandro Abate, Patrick Forré et al.ICML 2025
- Provably Efficient Exploration in Inverse Constrained Reinforcement LearningBo Yue, Jian Li, Guiliang LiuICML 2025
- Robustness in the Face of Partial Identifiability in Reward LearningFilippo Lazzati, Alberto Maria MetelliICLR 2026
Builds on10
- Quantifying Differences in Reward FunctionsAdam Gleave, Michael Dennis, Shane Legg, Stuart Russell et al.ICLR 2021 · 77 citations
- Identifiability in inverse reinforcement learningHaoyang Cao, Samuel N. Cohen, Lukasz SzpruchNeurIPS 2021 · 72 citations
- Avoiding Side Effects By Considering Future TasksVictoria Krakovna, Laurent Orseau, Richard Ngo, Miljan Martic et al.NeurIPS 2020 · 54 citations
- Reward Identification in Inverse Reinforcement LearningKuno Kim, Shivam Garg, Kirankumar Shiragur, Stefano ErmonICML 2021 · 43 citations
- Robust Inverse Reinforcement Learning under Transition Dynamics MismatchLuca Viano, Yu-Ting Huang, Parameswaran Kamalaruban, Adrian Weller et al.NeurIPS 2021 · 40 citations
Related papers
- Maximum Likelihood Constraint Inference for Inverse Reinforcement LearningDexter R. R. Scobee, S. Shankar SastryICLR 2020 · 74 citations
- Preference Learning for AI Alignment: a Causal PerspectiveKasia Kobalczyk, Mihaela van der SchaarICML 2025
- Inverse Reinforcement Learning with Explicit Policy EstimatesNavyata Sanghvi, Shinnosuke Usami, Mohit Sharma, Joachim Groeger et al.AAAI 2021 · 7 citations
- Inverse Reinforcement Learning in a Continuous State Space with Formal GuaranteesGregory Dexter, Kevin Bello, Jean HonorioNeurIPS 2021 · 9 citations
- Efficient Imitation under MisspecificationNicolas A. Espinosa Dice, Sanjiban Choudhury, Wen Sun, Gokul SwamyICLR 2025
