Off-Dynamics Reinforcement Learning: Training for Transfer with Domain Classifiers
Benjamin Eysenbach, Shreyas Chaudhari, Swapnil Asawa, Sergey Levine, Ruslan Salakhutdinov
Abstract
We propose a simple, practical, and intuitive approach for domain adaptation in reinforcement learning. Our approach stems from the idea that the agent's experience in the source domain should look similar to its experience in the target domain. Building off of a probabilistic view of RL, we achieve this goal by compensating for the difference in dynamics by modifying the reward function. This modified reward function is simple to estimate by learning auxiliary classifiers that distinguish source-domain transitions from target-domain transitions. Intuitively, the agent is penalized for transitions that would indicate that the agent is interacting with the source domain, rather than the target domain. Formally, we prove that applying our method in the source domain is guaranteed to obtain a near-optimal policy for the target domain, provided that the source and target domains satisfy a lightweight assumption. Our approach is applicable to domains with continuous states and actions and does not require learning an explicit model of the dynamics. On discrete and continuous control tasks, we illustrate the mechanics of our approach and demonstrate its scalability to high-dimensional tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers43
- When to Trust Your Simulator: Dynamics-Aware Hybrid Offline-and-Online Reinforcement LearningHaoyi Niu, Shubham Sharma, Yiwen Qiu, Ming Li et al.NeurIPS 2022 · 81 citations
- Mismatched No More: Joint Model-Policy Optimization for Model-Based RLBenjamin Eysenbach, Alexander Khazatsky, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 57 citations
- DARA: Dynamics-Aware Reward Augmentation in Offline Reinforcement LearningJinxin Liu, Hongyin Zhang, Donglin WangICLR 2022 · 47 citations
- Cross-Domain Policy Adaptation via Value-Guided Data FilteringKang Xu, Chenjia Bai, Xiaoteng Ma, Dong Wang et al.NeurIPS 2023 · 41 citations
- You Only Live Once: Single-Life Reinforcement LearningAnnie S. Chen, Archit Sharma, Sergey Levine, Chelsea FinnNeurIPS 2022 · 33 citations
Related papers
- Cross-Domain Policy Adaptation by Capturing Representation MismatchJiafei Lyu, Chenjia Bai, Jingwen Yang, Zongqing Lu et al.ICML 2024 · 30 citations
- Unsupervised Domain Adaptation with Dynamics-Aware Rewards in Reinforcement LearningJinxin Liu, Hao Shen, Donglin Wang, Yachen Kang et al.NeurIPS 2021 · 20 citations
- Off-Dynamics Reinforcement Learning via Domain Adaptation and Reward Augmented ImitationYihong Guo, Yixuan Wang, Yuanyuan Shi, Pan Xu et al.NeurIPS 2024 · 21 citations
- Online Prototype Alignment for Few-shot Policy TransferQi Yi, Rui Zhang, Shaohui Peng, Jiaming Guo et al.ICML 2023 · 5 citations
- Contrastive Representation for Data Filtering in Cross-Domain Offline Reinforcement LearningXiaoyu Wen, Chenjia Bai, Kang Xu, Xudong Yu et al.ICML 2024 · 13 citations
