Efficient Imitation under Misspecification
Nicolas A. Espinosa Dice, Sanjiban Choudhury, Wen Sun, Gokul Swamy
Abstract
We consider the problem of imitation learning under misspecification: settings where the learner is fundamentally unable to replicate expert behavior everywhere. This is often true in practice due to differences in observation space and action space expressiveness (e.g. perceptual or morphological differences between robots and humans). Given the learner must make some mistakes in the misspecified setting, interaction with the environment is fundamentally required to figure out which mistakes are particularly costly and lead to compounding errors. However, given the computational cost and safety concerns inherent in interaction, we would like to perform as little of it as possible while ensuring we have learned a strong policy. Accordingly, prior work has proposed a flavor of efficient inverse reinforcement learning algorithms that merely perform a computationally efficient local search procedure with strong guarantees in the realizable setting. We first prove that under a novel structural condition we term reward-agnostic policy completeness, these sorts of local-search based IRL algorithms are able to avoid compounding errors. We then consider the question of where we should perform local search in the first place, given the learner may not be able to "walk on a tightrope" as well as the expert in the misspecified setting. We prove that in the misspecified setting, it is beneficial to broaden the set of states on which local search is performed to include states reachable by good policies that the learner can actually play. We then experimentally explore a variety of sources of misspecification and how offline data can be used to effectively broaden where we perform local search from.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- A Smooth Sea Never Made a Skilled SAILOR: Robust Imitation via Learning to SearchArnav Kumar Jain, Vibhakar Mohta, Subin Kim, Atiksh Bhardwaj et al.NeurIPS 2025 · 27 citations
- Inverse Q-Learning Done Right: Offline Imitation Learning in Qπ-Realizable MDPsAntoine Moulin, Gergely Neu, Luca VianoNeurIPS 2025 · 6 citations
Builds on18
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- Of Moments and Matching: A Game-Theoretic Framework for Closing the Imitation GapGokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell, Steven WuICML 2021 · 90 citations
- Mitigating Covariate Shift in Imitation Learning via Offline Data With Partial CoverageJonathan D. Chang, Masatoshi Uehara, Dhruv Sreenivas, Rahul Kidambi et al.NeurIPS 2021 · 90 citations
- TRAIL: Near-Optimal Imitation Learning with Suboptimal DataMengjiao Yang, Sergey Levine, Ofir NachumICLR 2022 · 54 citations
- MobILE: Model-Based Imitation Learning From Observation AloneRahul Kidambi, Jonathan D. Chang, Wen SunNeurIPS 2021 · 51 citations
Related papers
- Hybrid Inverse Reinforcement LearningJuntao Ren, Gokul Swamy, Steven Wu, Drew Bagnell et al.ICML 2024 · 33 citations
- Offline Inverse RL: New Solution Concepts and Provably Efficient AlgorithmsFilippo Lazzati, Mirco Mutti, Alberto Maria MetelliICML 2024 · 8 citations
- Learning Shared Safety Constraints from Multi-task DemonstrationsKonwoo Kim, Gokul Swamy, Zuxin Liu, Ding Zhao et al.NeurIPS 2023 · 31 citations
- Non-Adversarial Inverse Reinforcement Learning via Successor Feature MatchingArnav Kumar Jain, Harley Wiltzer, Jesse Farebrother, Irina Rish et al.ICLR 2025
- Off-Policy Imitation Learning from ObservationsZhuangdi Zhu, Kaixiang Lin, Bo Dai, Jiayu ZhouNeurIPS 2020 · 102 citations
