Wasserstein Unsupervised Reinforcement Learning
Shuncheng He, Yuhang Jiang, Hongchang Zhang, Jianzhun Shao, Xiangyang Ji
摘要
Unsupervised reinforcement learning aims to train agents to learn a handful of policies or skills in environments without external reward. These pre-trained policies can accelerate learning when endowed with external reward, and can also be used as primitive options in hierarchical reinforcement learning. Conventional approaches of unsupervised skill discovery feed a latent variable to the agent and shed its empowerment on agent's behavior by mutual information (MI) maximization. However, the policies learned by MI-based methods cannot sufficiently explore the state space, despite they can be successfully identified from each other. Therefore we propose a new framework Wasserstein unsupervised reinforcement learning (WURL) where we directly maximize the distance of state distributions induced by different policies. Additionally, we overcome difficulties in simultaneously training N (N > 2) policies, and amortizing the overall reward to each step. Experiments show policies learned by our approach outperform MI-based methods on the metric of Wasserstein distance while keeping high discriminability. Furthermore, the agents trained by WURL can sufficiently explore the state space in mazes and MuJoCo tasks and the pre-trained policies can be applied to downstream tasks by hierarchical learning. Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- METRA: Scalable Unsupervised RL with Metric-Aware AbstractionSeohong Park, Oleh Rybkin, Sergey LevineICLR 2024 · 被引用 83 次
- Multi-Level Optimal Transport for Universal Cross-Tokenizer Knowledge Distillation on Language ModelsXiao Cui, Mo Zhu, Yulei Qin, Liang Xie 等AAAI 2025 · 被引用 31 次
- Challenging Common Assumptions in Convex Reinforcement LearningMirco Mutti, Riccardo De Santi, Piersilvio De Bartolomeis, Marcello RestelliNeurIPS 2022 · 被引用 31 次
- PEAC: Unsupervised Pre-training for Cross-Embodiment Reinforcement LearningChengyang Ying, Zhongkai Hao, Xinning Zhou, Xuezhou Xu 等NeurIPS 2024 · 被引用 14 次
- Global Reinforcement Learning : Beyond Linear and Convex Rewards via Submodular Semi-gradient MethodsRiccardo De Santi, Manish Prajapat, Andreas KrauseICML 2024 · 被引用 14 次
它引用的顶会 Paper7
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar 等ICLR 2020 · 被引用 475 次
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair 等ICML 2020 · 被引用 303 次
- Explore, Discover and Learn: Unsupervised Discovery of State-Covering SkillsVictor Campos, Alexander Trott, Caiming Xiong, Richard Socher 等ICML 2020 · 被引用 178 次
- Relative Variational Intrinsic ControlKate Baumli, David Warde-Farley, Steven Hansen, Volodymyr MnihAAAI 2021 · 被引用 45 次
- Learning to Score Behaviors for Guided Policy OptimizationAldo Pacchiano, Jack Parker-Holder, Yunhao Tang, Krzysztof Choromanski 等ICML 2020 · 被引用 42 次
相关 Paper
- Task Adaptation from Skills: Information Geometry, Disentanglement, and New Objectives for Unsupervised Reinforcement LearningYucheng Yang, Tianyi Zhou, Qiang He, Lei Han 等ICLR 2024 · 被引用 13 次
- Behavior Contrastive Learning for Unsupervised Skill DiscoveryRushuai Yang, Chenjia Bai, Hongyi Guo, Siyuan Li 等ICML 2023 · 被引用 34 次
- Skill Disentanglement in Reproducing Kernel Hilbert SpaceVedant Dave, Elmar RueckertAAAI 2025
- Lipschitz-constrained Unsupervised Skill DiscoverySeohong Park, Jongwook Choi, Jaekyeom Kim, Honglak Lee 等ICLR 2022 · 被引用 72 次
- The Information Geometry of Unsupervised Reinforcement LearningBenjamin Eysenbach, Ruslan Salakhutdinov, Sergey LevineICLR 2022 · 被引用 41 次
