Off-Dynamics Reinforcement Learning via Domain Adaptation and Reward Augmented Imitation
Yihong Guo, Yixuan Wang, Yuanyuan Shi, Pan Xu, Anqi Liu
Abstract
Training a policy in a source domain for deployment in the target domain under a dynamics shift can be challenging, often resulting in performance degradation. Previous work tackles this challenge by training on the source domain with modified rewards derived by matching distributions between the source and the target optimal trajectories. However, pure modified rewards only ensure the behavior of the learned policy in the source domain resembles trajectories produced by the target optimal policies, which does not guarantee optimal performance when the learned policy is actually deployed to the target domain. In this work, we propose to utilize imitation learning to transfer the policy learned from the reward modification to the target domain so that the new policy can generate the same trajectories in the target domain. Our approach, Domain Adaptation and Reward Augmented Imitation Learning (DARAIL), utilizes the reward modification for domain adaptation and follows the general framework of generative adversarial imitation learning from observation (GAIfO) by applying a reward augmented estimator for the policy optimization step. Theoretically, we present an error bound for our method under a mild assumption regarding the dynamics shift to justify the motivation of our method. Empirically, our method outperforms the pure modified reward method without imitation learning and also outperforms other baselines in benchmark off-dynamics environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 18f3b579-16c8-4b80-b7f3-83c8473ee095Cited by top-tier papers8
- MOBODY: Model-Based Off-Dynamics Offline Reinforcement LearningYihong Guo, Yu Yang, Pan Xu, Anqi LiuICLR 2026 · 10 citations
- Offline Guarded Safe Reinforcement Learning for Medical Treatment Optimization StrategiesRunze Yan, Xun Shen, Akifumi Wachi, Sebastien Gros et al.NeurIPS 2025 · 7 citations
- Linear Mixture Distributionally Robust Markov Decision ProcessesZhishuai Liu, Pan XuNeurIPS 2025 · 6 citations
- Dynamics Are Learned, Not Told: Semi-Supervised Discovery of Latent Dynamics Geometries For Zero-Shot Policy AdaptationZhiming Xu, Weitao Zhou, Xianghui Pan, Nanshan Deng et al.ICML 2026
- Transport or Discard: Robust Unbalanced Optimal Transport for Cross-Domain Policy AdaptationWenyu Chen, Yujia Zhang, Wei Guo, Linli Ma et al.ICML 2026
Builds on11
- Error Bounds of Imitating Policies and EnvironmentsTian Xu, Ziniu Li, Yang YuNeurIPS 2020 · 141 citations
- Doubly robust off-policy evaluation with shrinkageYi Su, Maria Dimakopoulou, Akshay Krishnamurthy, Miroslav DudíkICML 2020 · 128 citations
- Off-Dynamics Reinforcement Learning: Training for Transfer with Domain ClassifiersBenjamin Eysenbach, Shreyas Chaudhari, Swapnil Asawa, Sergey Levine et al.ICLR 2021 · 120 citations
- Domain Adaptive Imitation LearningKuno Kim, Yihong Gu, Jiaming Song, Shengjia Zhao et al.ICML 2020 · 86 citations
- Domain Adaptation In Reinforcement Learning Via Latent Unified State RepresentationJinwei Xing, Takashi Nagata, Kexin Chen, Xinyun Zou et al.AAAI 2021 · 65 citations
Related papers
- An Imitation from Observation Approach to Transfer Learning with Dynamics MismatchSiddharth Desai, Ishan Durugkar, Haresh Karnan, Garrett Warnell et al.NeurIPS 2020 · 60 citations
- DARA: Dynamics-Aware Reward Augmentation in Offline Reinforcement LearningJinxin Liu, Hongyin Zhang, Donglin WangICLR 2022 · 47 citations
- Diffusion-Reward Adversarial Imitation LearningChun-Mao Lai, Hsiang-Chun Wang, Ping-Chun Hsieh, Yu-Chiang Frank Wang et al.NeurIPS 2024 · 28 citations
- Doubly Robust Augmented Transfer for Meta-Reinforcement LearningYuankun Jiang, Nuowen Kan, Chenglin Li, Wenrui Dai et al.NeurIPS 2023 · 3 citations
- Cross-Domain Policy Adaptation by Capturing Representation MismatchJiafei Lyu, Chenjia Bai, Jingwen Yang, Zongqing Lu et al.ICML 2024 · 30 citations
