Mitigating Covariate Shift in Imitation Learning via Offline Data With Partial Coverage
Jonathan D. Chang, Masatoshi Uehara, Dhruv Sreenivas, Rahul Kidambi, Wen Sun
Abstract
This paper studies offline Imitation Learning (IL) where an agent learns to imitate an expert demonstrator without additional online environment interactions. Instead, the learner is presented with a static offline dataset of state-action-next state transition triples from a potentially less proficient behavior policy. We introduce Model-based IL from Offline data (MILO): an algorithmic framework that utilizes the static dataset to solve the offline IL problem efficiently both in theory and in practice. In theory, even if the behavior policy is highly sub-optimal compared to the expert, we show that as long as the data from the behavior policy provides sufficient coverage on the expert state-action traces (and with no necessity for a global coverage over the entire state-action space), MILO can provably combat the covariate shift issue in IL. Complementing our theory results, we also demonstrate that a practical implementation of our approach mitigates covariate shift on benchmark MuJoCo continuous control tasks. We demonstrate that with behavior policies whose performances are less than half of that of the expert, MILO still successfully imitates with an extremely low number of expert state-action pairs while traditional offline IL methods such as behavior cloning (BC) fail completely. Source code is provided at https://github.com/jdchang1/milo . * Equal contribution † Work done outside Amazon 35th Conference on Neural Information Processing Systems (NeurIPS 2021), virtual.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c545d8a-7a9f-4777-bc7b-be5c226863fcCited by top-tier papers41
- Chain of Thought Imitation with Procedure CloningMengjiao Yang, Dale Schuurmans, Pieter Abbeel, Ofir NachumNeurIPS 2022 · 53 citations
- Leveraging Offline Data in Online Reinforcement LearningAndrew Wagenmaker, Aldo PacchianoICML 2023 · 47 citations
- When Demonstrations meet Generative World Models: A Maximum Likelihood Framework for Offline Inverse Reinforcement LearningSiliang Zeng, Chenliang Li, Alfredo García, Mingyi HongNeurIPS 2023 · 33 citations
- Latent Policy Barrier: Learning Robust Visuomotor Policies by Staying In-DistributionZhanyi Sun, Shuran SongNeurIPS 2025 · 27 citations
- Proximal Point Imitation LearningLuca Viano, Angeliki Kamoutsi, Gergely Neu, Igor Krawczuk et al.NeurIPS 2022 · 27 citations
Builds on27
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
- COMBO: Conservative Offline Model-Based Policy OptimizationTianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran et al.NeurIPS 2021 · 549 citations
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 419 citations
Related papers
- Discriminator-Weighted Offline Imitation Learning from Suboptimal DemonstrationsHaoran Xu, Xianyuan Zhan, Honglei Yin, Huiling QinICML 2022 · 105 citations
- Curriculum Offline Imitating LearningMinghuan Liu, Hanye Zhao, Zhengyu Yang, Jian Shen et al.NeurIPS 2021 · 5 citations
- Offline Imitation Learning with Model-based Reverse AugmentationJie-Jing Shao, Hao-Sen Shi, Lan-Zhe Guo, Yu-Feng LiKDD 2024 · 5 citations
- Mitigating Covariate Shift in Behavioral Cloning via Robust Stationary Distribution CorrectionSeokin Seo, Byung-Jun Lee, Jongmin Lee, HyeongJoo Hwang et al.NeurIPS 2024 · 17 citations
- DemoDICE: Offline Imitation Learning with Supplementary Imperfect DemonstrationsGeon-Hyeong Kim, Seokin Seo, Jongmin Lee, Wonseok Jeon et al.ICLR 2022 · 111 citations
