Off-Dynamics Reinforcement Learning: Training for Transfer with Domain Classifiers
Benjamin Eysenbach, Shreyas Chaudhari, Swapnil Asawa, Sergey Levine, Ruslan Salakhutdinov
摘要
We propose a simple, practical, and intuitive approach for domain adaptation in reinforcement learning. Our approach stems from the idea that the agent's experience in the source domain should look similar to its experience in the target domain. Building off of a probabilistic view of RL, we achieve this goal by compensating for the difference in dynamics by modifying the reward function. This modified reward function is simple to estimate by learning auxiliary classifiers that distinguish source-domain transitions from target-domain transitions. Intuitively, the agent is penalized for transitions that would indicate that the agent is interacting with the source domain, rather than the target domain. Formally, we prove that applying our method in the source domain is guaranteed to obtain a near-optimal policy for the target domain, provided that the source and target domains satisfy a lightweight assumption. Our approach is applicable to domains with continuous states and actions and does not require learning an explicit model of the dynamics. On discrete and continuous control tasks, we illustrate the mechanics of our approach and demonstrate its scalability to high-dimensional tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper43
- When to Trust Your Simulator: Dynamics-Aware Hybrid Offline-and-Online Reinforcement LearningHaoyi Niu, Shubham Sharma, Yiwen Qiu, Ming Li 等NeurIPS 2022 · 被引用 81 次
- Mismatched No More: Joint Model-Policy Optimization for Model-Based RLBenjamin Eysenbach, Alexander Khazatsky, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 被引用 57 次
- DARA: Dynamics-Aware Reward Augmentation in Offline Reinforcement LearningJinxin Liu, Hongyin Zhang, Donglin WangICLR 2022 · 被引用 47 次
- Cross-Domain Policy Adaptation via Value-Guided Data FilteringKang Xu, Chenjia Bai, Xiaoteng Ma, Dong Wang 等NeurIPS 2023 · 被引用 41 次
- You Only Live Once: Single-Life Reinforcement LearningAnnie S. Chen, Archit Sharma, Sergey Levine, Chelsea FinnNeurIPS 2022 · 被引用 33 次
相关 Paper
- Cross-Domain Policy Adaptation by Capturing Representation MismatchJiafei Lyu, Chenjia Bai, Jingwen Yang, Zongqing Lu 等ICML 2024 · 被引用 30 次
- Unsupervised Domain Adaptation with Dynamics-Aware Rewards in Reinforcement LearningJinxin Liu, Hao Shen, Donglin Wang, Yachen Kang 等NeurIPS 2021 · 被引用 20 次
- Off-Dynamics Reinforcement Learning via Domain Adaptation and Reward Augmented ImitationYihong Guo, Yixuan Wang, Yuanyuan Shi, Pan Xu 等NeurIPS 2024 · 被引用 21 次
- Online Prototype Alignment for Few-shot Policy TransferQi Yi, Rui Zhang, Shaohui Peng, Jiaming Guo 等ICML 2023 · 被引用 5 次
- Contrastive Representation for Data Filtering in Cross-Domain Offline Reinforcement LearningXiaoyu Wen, Chenjia Bai, Kang Xu, Xudong Yu 等ICML 2024 · 被引用 13 次
