Robust Imitation of a Few Demonstrations with a Backwards Model
Jung Yeon Park, Lawson L. S. Wong
Abstract
Behavior cloning of expert demonstrations can speed up learning optimal policies in a more sample-efficient way over reinforcement learning. However, the policy cannot extrapolate well to unseen states outside of the demonstration data, creating covariate shift (agent drifting away from demonstrations) and compounding errors. In this work, we tackle this issue by extending the region of attraction around the demonstrations so that the agent can learn how to get back onto the demonstrated trajectories if it veers off-course. We train a generative backwards dynamics model and generate short imagined trajectories from states in the demonstrations. By imitating both demonstrations and these model rollouts, the agent learns the demonstrated paths and how to get back onto these paths. With optimal or near-optimal demonstrations, the learned policy will be both optimal and robust to deviations, with a wider region of attraction. On continuous control domains, we evaluate the robustness when starting from different initial states unseen in the demonstration data. While both our method and other imitation learning baselines can successfully solve the tasks for initial states in the training distribution, our method exhibits considerably more robustness to different initial states.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- CCIL: Continuity-Based Data Augmentation for Corrective Imitation LearningLiyiming Ke, Yunchu Zhang, Abhay Deshpande, Siddhartha S. Srinivasa et al.ICLR 2024 · 33 citations
- Latent Policy Barrier: Learning Robust Visuomotor Policies by Staying In-DistributionZhanyi Sun, Shuran SongNeurIPS 2025 · 27 citations
- From Past to Future: Rethinking Eligibility TracesDhawal Gupta, Scott M. Jordan, Shreyas Chaudhari, Bo Liu et al.AAAI 2024 · 5 citations
Builds on12
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- How Neural Networks Extrapolate: From Feedforward to Graph Neural NetworksKeyulu Xu, Mozhi Zhang, Jingling Li, Simon Shaolei Du et al.ICLR 2021 · 364 citations
- SQIL: Imitation Learning via Reinforcement Learning with Sparse RewardsSiddharth Reddy, Anca D. Dragan, Sergey LevineICLR 2020 · 299 citations
- Context-aware Dynamics Model for Generalization in Model-Based Reinforcement LearningKimin Lee, Younggyo Seo, Seunghyun Lee, Honglak Lee et al.ICML 2020 · 158 citations
- Offline Reinforcement Learning with Reverse Model-based ImaginationJianhao Wang, Wenzhe Li, Haozhe Jiang, Guangxiang Zhu et al.NeurIPS 2021 · 74 citations
Related papers
- Contractive Dynamical Imitation Policies for Efficient Out-of-Sample RecoveryAmin Abyaneh, Mahrokh Ghoddousi Boroujeni, Hsiu-Chin Lin, Giancarlo Ferrari-TrecateICLR 2025
- Learning Self-Correctable Policies and Value Functions from Demonstrations with Negative SamplingYuping Luo, Huazhe Xu, Tengyu MaICLR 2020 · 14 citations
- Offline Imitation Learning with Model-based Reverse AugmentationJie-Jing Shao, Hao-Sen Shi, Lan-Zhe Guo, Yu-Feng LiKDD 2024 · 5 citations
- Disagreement-Regularized Imitation LearningKianté Brantley, Wen Sun, Mikael HenaffICLR 2020 · 112 citations
- Behavioral Cloning from Noisy DemonstrationsFumihiro Sasaki, Ryota YamashinaICLR 2021 · 94 citations
