Visual Transfer For Reinforcement Learning Via Wasserstein Domain Confusion
Josh Roy, George Dimitri Konidaris
Abstract
We introduce Wasserstein Adversarial Proximal Policy Optimization (WAPPO), a novel algorithm for visual transfer in Reinforcement Learning that explicitly learns to align the distributions of extracted features between a source and target task. WAPPO approximates and minimizes the Wasserstein-1 distance between the distributions of features from source and target domains via a novel Wasserstein Confusion objective. WAPPO outperforms the prior state-of-the-art in visual transfer and successfully transfers policies across Visual Cartpole and two instantiations of 16 OpenAI Procgen environments. * joshnroy.github.io Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0de3088b-103e-4c4c-bdac-46bfcc2a3135Cited by top-tier papers5
- Automatic Data Augmentation for Generalization in Reinforcement LearningRoberta Raileanu, Maxwell Goldstein, Denis Yarats, Ilya Kostrikov et al.NeurIPS 2021 · 143 citations
- Decoupling Value and Policy for Generalization in Reinforcement LearningRoberta Raileanu, Rob FergusICML 2021 · 116 citations
- On the Importance of Exploration for Generalization in Reinforcement LearningYiding Jiang, J. Zico Kolter, Roberta RaileanuNeurIPS 2023 · 48 citations
- The Generalization Gap in Offline Reinforcement LearningIshita Mediratta, Qingfei You, Minqi Jiang, Roberta RaileanuICLR 2024 · 24 citations
- Know Thyself: Transferable Visual Control Policies Through Robot-AwarenessEdward S. Hu, Kun Huang, Oleh Rybkin, Dinesh JayaramanICLR 2022 · 2 citations
Builds on1
Related papers
- Prompt-based Visual Alignment for Zero-shot Policy TransferHaihan Gao, Rui Zhang, Qi Yi, Hantao Yao et al.ICML 2024 · 1 citation
- Online Prototype Alignment for Few-shot Policy TransferQi Yi, Rui Zhang, Shaohui Peng, Jiaming Guo et al.ICML 2023 · 5 citations
- Domain Adaptive Imitation Learning with Visual ObservationSungho Choi, Seungyul Han, Woojun Kim, Jongseong Chae et al.NeurIPS 2023 · 15 citations
- Relative Policy-Transition Optimization for Fast Policy TransferJiawei Xu, Cheng Zhou, Yizheng Zhang, Baoxiang Wang et al.AAAI 2024
- Trust Region Policy Optimization with Optimal Transport Discrepancies: Duality and Algorithm for Continuous ActionsAntonio Terpin, Nicolas Lanzetti, Batuhan Yardim, Florian Dörfler et al.NeurIPS 2022 · 14 citations
