Offline Imitation Learning with a Misspecified Simulator
Shengyi Jiang, Jing-Cheng Pang, Yang Yu
Abstract
In real-world decision-making tasks, learning an optimal policy without a trialand-error process is an appealing challenge. When expert demonstrations are available, imitation learning that mimics expert actions can learn a good policy efficiently. Learning in simulators is another commonly adopted approach to avoid real-world trials-and-errors. However, neither sufficient expert demonstrations nor high-fidelity simulators are easy to obtain. In this work, we investigate policy learning in the condition of a few expert demonstrations and a simulator with misspecified dynamics. Under a mild assumption that local states shall still be partially aligned under a dynamics mismatch, we propose imitation learning with horizon-adaptive inverse dynamics (HIDIL) that matches the simulator states with expert states in a H-step horizon and accurately recovers actions based on inverse dynamics policies. In the real environment, HIDIL can effectively derive adapted actions from the matched states. Experiments are conducted in four MuJoCo locomotion environments with modified friction, gravity, and density configurations. Experiment results show that HIDIL achieves significant improvement in terms of performance and stability in all of the real environments, compared with imitation learning methods and transferring methods in reinforcement learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 21a1cb3c-77de-4e35-878a-ea99efbf7e2eCited by top-tier papers10
- State Regularized Policy Optimization on Data with Dynamics ShiftZhenghai Xue, Qingpeng Cai, Shuchang Liu, Dong Zheng et al.NeurIPS 2023 · 30 citations
- CEIL: Generalized Contextual Imitation LearningJinxin Liu, Li He, Yachen Kang, Zifeng Zhuang et al.NeurIPS 2023 · 23 citations
- Off-Dynamics Reinforcement Learning via Domain Adaptation and Reward Augmented ImitationYihong Guo, Yixuan Wang, Yuanyuan Shi, Pan Xu et al.NeurIPS 2024 · 21 citations
- Mitigating Covariate Shift in Behavioral Cloning via Robust Stationary Distribution CorrectionSeokin Seo, Byung-Jun Lee, Jongmin Lee, HyeongJoo Hwang et al.NeurIPS 2024 · 17 citations
- Plan Your Target and Learn Your Skills: Transferable State-Only Imitation Learning via Decoupled Policy OptimizationMinghuan Liu, Zhengbang Zhu, Yuzheng Zhuang, Weinan Zhang et al.ICML 2022 · 13 citations
Builds on3
- Disagreement-Regularized Imitation LearningKianté Brantley, Wen Sun, Mikael HenaffICLR 2020 · 112 citations
- State Alignment-based Imitation LearningFangchen Liu, Zhan Ling, Tongzhou Mu, Hao SuICLR 2020 · 103 citations
- State-only Imitation with Transition Dynamics MismatchTanmay Gangwani, Jian PengICLR 2020 · 56 citations
Related papers
- Robust Visual Imitation Learning with Inverse Dynamics RepresentationsSiyuan Li, Xun Wang, Rongchang Zuo, Kewu Sun et al.AAAI 2024 · 8 citations
- Mimicking Better by Matching the Approximate Action DistributionJoão A. Cândido Ramos, Lionel Blondé, Naoya Takeishi, Alexandros KalousisICML 2024 · 4 citations
- Imitation Learning from Observations under Transition Model DisparityTanmay Gangwani, Yuan Zhou, Jian PengICLR 2022 · 15 citations
- Unlabeled Imperfect Demonstrations in Adversarial Imitation LearningYunke Wang, Bo Du, Chang XuAAAI 2023 · 11 citations
- Seeing Differently, Acting Similarly: Heterogeneously Observable Imitation LearningXin-Qiang Cai, Yao-Xiang Ding, Zi-Xuan Chen, Yuan Jiang et al.ICLR 2023 · 2 citations
