The Information Geometry of Unsupervised Reinforcement Learning
Benjamin Eysenbach, Ruslan Salakhutdinov, Sergey Levine
Abstract
How can a reinforcement learning (RL) agent prepare to solve downstream tasks if those tasks are not known a priori? One approach is unsupervised skill discovery, a class of algorithms that learn a set of policies without access to a reward function. Such algorithms bear a close resemblance to representation learning algorithms (e.g., contrastive learning) in supervised learning, in that both are pretraining algorithms that maximize some approximation to a mutual information objective. While prior work has shown that the set of skills learned by such methods can accelerate downstream RL tasks, prior work offers little analysis into whether these skill learning algorithms are optimal, or even what notion of optimality would be appropriate to apply to them. In this work, we show that unsupervised skill discovery algorithms based on mutual information maximization do not learn skills that are optimal for every possible reward function. However, we show that the distribution over skills provides an optimal initialization minimizing regret against adversarially-chosen reward functions, assuming a certain type of adaptation procedure. Our analysis also provides a geometric perspective on these skill learning methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 55116606-d7f2-4448-8dca-b05f81cf1e7dCited by top-tier papers26
- Optimistic Linear Support and Successor Features as a Basis for Optimal Policy TransferLucas Nunes Alegre, Ana L. C. Bazzan, Bruno C. da SilvaICML 2022 · 36 citations
- Acquiring Diverse Skills using Curriculum Reinforcement Learning with Mixture of ExpertsOnur Celik, Aleksandar Taranovic, Gerhard NeumannICML 2024 · 19 citations
- Learning to Assist Humans without Inferring RewardsVivek Myers, Evan Ellis, Sergey Levine, Benjamin Eysenbach et al.NeurIPS 2024 · 16 citations
- A Mixture Of Surprises for Unsupervised Reinforcement LearningAndrew Zhao, Matthieu Gaetan Lin, Yangguang Li, Yong-Jin Liu et al.NeurIPS 2022 · 15 citations
- PEAC: Unsupervised Pre-training for Cross-Embodiment Reinforcement LearningChengyang Ying, Zhongkai Hao, Xinning Zhou, Xuezhou Xu et al.NeurIPS 2024 · 14 citations
Builds on11
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel et al.ICML 2020 · 489 citations
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar et al.ICLR 2020 · 475 citations
- Understanding the Limitations of Variational Mutual Information EstimatorsJiaming Song, Stefano ErmonICLR 2020 · 243 citations
- Explore, Discover and Learn: Unsupervised Discovery of State-Covering SkillsVictor Campos, Alexander Trott, Caiming Xiong, Richard Socher et al.ICML 2020 · 178 citations
Related papers
- Behavior Contrastive Learning for Unsupervised Skill DiscoveryRushuai Yang, Chenjia Bai, Hongyi Guo, Siyuan Li et al.ICML 2023 · 34 citations
- Task Adaptation from Skills: Information Geometry, Disentanglement, and New Objectives for Unsupervised Reinforcement LearningYucheng Yang, Tianyi Zhou, Qiang He, Lei Han et al.ICLR 2024 · 13 citations
- Efficient Skill Discovery via Regret-Aware OptimizationHe Zhang, Ming Zhou, Shaopeng Zhai, Ying Sun et al.ICML 2025
- Wasserstein Unsupervised Reinforcement LearningShuncheng He, Yuhang Jiang, Hongchang Zhang, Jianzhun Shao et al.AAAI 2022 · 30 citations
- Unsupervised Reinforcement Learning with Contrastive Intrinsic ControlMichael Laskin, Hao Liu, Xue Bin Peng, Denis Yarats et al.NeurIPS 2022 · 62 citations
