Variational Curriculum Reinforcement Learning for Unsupervised Discovery of Skills
Seongun Kim, Kyowoon Lee, Jaesik Choi
摘要
Mutual information-based reinforcement learning (RL) has been proposed as a promising framework for retrieving complex skills autonomously without a task-oriented reward function through mutual information (MI) maximization or variational empowerment. However, learning complex skills is still challenging, due to the fact that the order of training skills can largely affect sample efficiency. Inspired by this, we recast variational empowerment as curriculum learning in goal-conditioned RL with an intrinsic reward function, which we name Variational Curriculum RL (VCRL). From this perspective, we propose a novel approach to unsupervised skill discovery based on information theory, called Value Uncertainty Variational Curriculum (VUVC). We prove that, under regularity conditions, VUVC accelerates the increase of entropy in the visited states compared to the uniform curriculum. We validate the effectiveness of our approach on complex navigation and robotic manipulation tasks in terms of sample efficiency and state coverage speed. We also demonstrate that the skills discovered by our method successfully complete a real-world robot navigation task in a zero-shot setup and that incorporating these skills with a global planner further increases the performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- METRA: Scalable Unsupervised RL with Metric-Aware AbstractionSeohong Park, Oleh Rybkin, Sergey LevineICLR 2024 · 被引用 83 次
- Learning Multimodal Behaviors from Scratch with Diffusion Policy GradientSteven Li, Rickmer Krohn, Tao Chen, Anurag Ajay 等NeurIPS 2024 · 被引用 61 次
- How Far Can Unsupervised RLVR Scale LLM Training?Bingxiang He, Yuxin Zuo, Zeyuan Liu, Shangziqi Zhao 等ICLR 2026 · 被引用 31 次
- Refining Diffusion Planner for Reliable Behavior Synthesis by Automatic Detection of Infeasible PlansKyowoon Lee, Seongun Kim, Jaesik ChoiNeurIPS 2023 · 被引用 31 次
- State-Covering Trajectory Stitching for Diffusion PlannersKyowoon Lee, Jaesik ChoiNeurIPS 2025 · 被引用 17 次
它引用的顶会 Paper16
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel 等ICML 2020 · 被引用 489 次
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar 等ICLR 2020 · 被引用 475 次
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair 等ICML 2020 · 被引用 303 次
- Reinforcement Learning with Prototypical RepresentationsDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICML 2021 · 被引用 262 次
- Behavior From the Void: Unsupervised Active Pre-TrainingHao Liu, Pieter AbbeelNeurIPS 2021 · 被引用 258 次
相关 Paper
- Variational Empowerment as Representation Learning for Goal-Conditioned Reinforcement LearningJongwook Choi, Archit Sharma, Honglak Lee, Sergey Levine 等ICML 2021 · 被引用 41 次
- Behavior Contrastive Learning for Unsupervised Skill DiscoveryRushuai Yang, Chenjia Bai, Hongyi Guo, Siyuan Li 等ICML 2023 · 被引用 34 次
- Information Prioritization through Empowerment in Visual Model-based RLHomanga Bharadhwaj, Mohammad Babaeizadeh, Dumitru Erhan, Sergey LevineICLR 2022 · 被引用 35 次
- Explore, Discover and Learn: Unsupervised Discovery of State-Covering SkillsVictor Campos, Alexander Trott, Caiming Xiong, Richard Socher 等ICML 2020 · 被引用 178 次
- Constrained Ensemble Exploration for Unsupervised Skill DiscoveryChenjia Bai, Rushuai Yang, Qiaosheng Zhang, Kang Xu 等ICML 2024 · 被引用 9 次
