An Imitation from Observation Approach to Transfer Learning with Dynamics Mismatch
Siddharth Desai, Ishan Durugkar, Haresh Karnan, Garrett Warnell, Josiah Hanna, Peter Stone
Abstract
We examine the problem of transferring a policy learned in a source environment to a target environment with different dynamics, particularly in the case where it is critical to reduce the amount of interaction with the target environment during learning. This problem is particularly important in sim-to-real transfer because simulators inevitably model real-world dynamics imperfectly. In this paper, we show that one existing solution to this transfer problem - grounded action transformation - is closely related to the problem of imitation from observation (IfO): learning behaviors that mimic the observations of behavior demonstrations. After establishing this relationship, we hypothesize that recent state-of-the-art approaches from the IfO literature can be effectively repurposed for grounded transfer learning.To validate our hypothesis we derive a new algorithm - generative adversarial reinforced action transformation (GARAT) - based on adversarial imitation from observation techniques. We run experiments in several domains with mismatched dynamics, and find that agents trained with GARAT achieve higher returns in the target environment compared to existing black-box transfer methods
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ccac5424-31b6-4077-a92c-c6bd2634c58dCited by top-tier papers18
- Causal Navigation by Continuous-time Neural NetworksCharles Vorbach, Ramin M. Hasani, Alexander Amini, Mathias Lechner et al.NeurIPS 2021 · 64 citations
- MobILE: Model-Based Imitation Learning From Observation AloneRahul Kidambi, Jonathan D. Chang, Wen SunNeurIPS 2021 · 51 citations
- Cross-Domain Policy Adaptation via Value-Guided Data FilteringKang Xu, Chenjia Bai, Xiaoteng Ma, Dong Wang et al.NeurIPS 2023 · 41 citations
- ASID: Active Exploration for System Identification in Robotic ManipulationMarius Memmel, Andrew Wagenmaker, Chuning Zhu, Dieter Fox et al.ICLR 2024 · 36 citations
- Cross-Domain Policy Adaptation by Capturing Representation MismatchJiafei Lyu, Chenjia Bai, Jingwen Yang, Zongqing Lu et al.ICML 2024 · 30 citations
Builds on1
Related papers
- Off-Dynamics Reinforcement Learning via Domain Adaptation and Reward Augmented ImitationYihong Guo, Yixuan Wang, Yuanyuan Shi, Pan Xu et al.NeurIPS 2024 · 21 citations
- Diffusion Imitation from ObservationBo-Ruei Huang, Chun-Kai Yang, Chun-Mao Lai, Dai-Jie Wu et al.NeurIPS 2024 · 15 citations
- Plan Your Target and Learn Your Skills: Transferable State-Only Imitation Learning via Decoupled Policy OptimizationMinghuan Liu, Zhengbang Zhu, Yuzheng Zhuang, Weinan Zhang et al.ICML 2022 · 13 citations
- Policy Regularization on Globally Accessible States in Cross-Dynamics Reinforcement LearningZhenghai Xue, Lang Feng, Jiacheng Xu, Kang Kang et al.ICML 2025
- Imitation Learning from Observations under Transition Model DisparityTanmay Gangwani, Yuan Zhou, Jian PengICLR 2022 · 15 citations
