Offline Imitation Learning with a Misspecified Simulator
Shengyi Jiang, Jing-Cheng Pang, Yang Yu
摘要
In real-world decision-making tasks, learning an optimal policy without a trialand-error process is an appealing challenge. When expert demonstrations are available, imitation learning that mimics expert actions can learn a good policy efficiently. Learning in simulators is another commonly adopted approach to avoid real-world trials-and-errors. However, neither sufficient expert demonstrations nor high-fidelity simulators are easy to obtain. In this work, we investigate policy learning in the condition of a few expert demonstrations and a simulator with misspecified dynamics. Under a mild assumption that local states shall still be partially aligned under a dynamics mismatch, we propose imitation learning with horizon-adaptive inverse dynamics (HIDIL) that matches the simulator states with expert states in a H-step horizon and accurately recovers actions based on inverse dynamics policies. In the real environment, HIDIL can effectively derive adapted actions from the matched states. Experiments are conducted in four MuJoCo locomotion environments with modified friction, gravity, and density configurations. Experiment results show that HIDIL achieves significant improvement in terms of performance and stability in all of the real environments, compared with imitation learning methods and transferring methods in reinforcement learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- State Regularized Policy Optimization on Data with Dynamics ShiftZhenghai Xue, Qingpeng Cai, Shuchang Liu, Dong Zheng 等NeurIPS 2023 · 被引用 30 次
- CEIL: Generalized Contextual Imitation LearningJinxin Liu, Li He, Yachen Kang, Zifeng Zhuang 等NeurIPS 2023 · 被引用 23 次
- Off-Dynamics Reinforcement Learning via Domain Adaptation and Reward Augmented ImitationYihong Guo, Yixuan Wang, Yuanyuan Shi, Pan Xu 等NeurIPS 2024 · 被引用 21 次
- Mitigating Covariate Shift in Behavioral Cloning via Robust Stationary Distribution CorrectionSeokin Seo, Byung-Jun Lee, Jongmin Lee, HyeongJoo Hwang 等NeurIPS 2024 · 被引用 17 次
- Plan Your Target and Learn Your Skills: Transferable State-Only Imitation Learning via Decoupled Policy OptimizationMinghuan Liu, Zhengbang Zhu, Yuzheng Zhuang, Weinan Zhang 等ICML 2022 · 被引用 13 次
它引用的顶会 Paper3
- Disagreement-Regularized Imitation LearningKianté Brantley, Wen Sun, Mikael HenaffICLR 2020 · 被引用 112 次
- State Alignment-based Imitation LearningFangchen Liu, Zhan Ling, Tongzhou Mu, Hao SuICLR 2020 · 被引用 103 次
- State-only Imitation with Transition Dynamics MismatchTanmay Gangwani, Jian PengICLR 2020 · 被引用 56 次
相关 Paper
- Robust Visual Imitation Learning with Inverse Dynamics RepresentationsSiyuan Li, Xun Wang, Rongchang Zuo, Kewu Sun 等AAAI 2024 · 被引用 8 次
- Mimicking Better by Matching the Approximate Action DistributionJoão A. Cândido Ramos, Lionel Blondé, Naoya Takeishi, Alexandros KalousisICML 2024 · 被引用 4 次
- Imitation Learning from Observations under Transition Model DisparityTanmay Gangwani, Yuan Zhou, Jian PengICLR 2022 · 被引用 15 次
- Unlabeled Imperfect Demonstrations in Adversarial Imitation LearningYunke Wang, Bo Du, Chang XuAAAI 2023 · 被引用 11 次
- Seeing Differently, Acting Similarly: Heterogeneously Observable Imitation LearningXin-Qiang Cai, Yao-Xiang Ding, Zi-Xuan Chen, Yuan Jiang 等ICLR 2023 · 被引用 2 次
