Learning more skills through optimistic exploration
DJ Strouse, Kate Baumli, David Warde-Farley, Volodymyr Mnih, Steven Stenberg Hansen
Abstract
Unsupervised skill learning objectives (Eysenbach et al., 2019; Gregor et al., 2016) allow agents to learn rich repertoires of behavior in the absence of extrinsic rewards. They work by simultaneously training a policy to produce distinguishable latent-conditioned trajectories, and a discriminator to evaluate distinguishability by trying to infer latents from trajectories. The hope is for the agent to explore and master the environment by encouraging each skill (latent) to reliably reach different states. However, an inherent exploration problem lingers: when a novel state is actually encountered, the discriminator will necessarily not have seen enough training data to produce accurate and confident skill classifications, leading to low intrinsic reward for the agent and effective penalization of the sort of exploration needed to actually maximize the objective. To combat this inherent pessimism towards exploration, we derive an information gain auxiliary objective that involves training an ensemble of discriminators and rewarding the policy for their disagreement. Our objective directly estimates the epistemic uncertainty that comes from the discriminator not having seen enough training examples, thus providing an intrinsic reward more tailored to the true objective compared to pseudocount-based methods (Burda et al., 2019) . We call this exploration bonus discriminator disagreement intrinsic reward, or DISDAIN. We demonstrate empirically that DISDAIN improves skill learning both in a tabular grid world (Four Rooms) and the 57 games of the Atari Suite (from pixels). Thus, we encourage researchers to treat pessimism with DISDAIN.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 60c832f2-a6b5-4c21-a078-d2b38a50c2feCited by top-tier papers27
- Deep Hierarchical Planning from PixelsDanijar Hafner, Kuang-Huei Lee, Ian Fischer, Pieter AbbeelNeurIPS 2022 · 153 citations
- Planning Goals for ExplorationEdward S. Hu, Richard Chang, Oleh Rybkin, Dinesh JayaramanICLR 2023 · 152 citations
- Semantic Exploration from Language Abstractions and Pretrained RepresentationsAllison C. Tam, Neil C. Rabinowitz, Andrew K. Lampinen, Nicholas A. Roy et al.NeurIPS 2022 · 85 citations
- METRA: Scalable Unsupervised RL with Metric-Aware AbstractionSeohong Park, Oleh Rybkin, Sergey LevineICLR 2024 · 83 citations
- Unsupervised Reinforcement Learning with Contrastive Intrinsic ControlMichael Laskin, Hao Liu, Xue Bin Peng, Denis Yarats et al.NeurIPS 2022 · 62 citations
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel et al.ICML 2020 · 489 citations
Related papers
- Learning to Discover Skills through GuidanceHyunseung Kim, Byungkun Lee, Hojoon Lee, Dongyoon Hwang et al.NeurIPS 2023 · 14 citations
- Skill Disentanglement in Reproducing Kernel Hilbert SpaceVedant Dave, Elmar RueckertAAAI 2025
- Behavior Contrastive Learning for Unsupervised Skill DiscoveryRushuai Yang, Chenjia Bai, Hongyi Guo, Siyuan Li et al.ICML 2023 · 34 citations
- Relative Variational Intrinsic ControlKate Baumli, David Warde-Farley, Steven Hansen, Volodymyr MnihAAAI 2021 · 45 citations
- SkiLD: Unsupervised Skill Discovery Guided by Factor InteractionsZizhao Wang, Jiaheng Hu, Caleb Chuck, Stephen Chen et al.NeurIPS 2024 · 15 citations
