Mastering the Unsupervised Reinforcement Learning Benchmark from Pixels
Sai Rajeswar, Pietro Mazzaglia, Tim Verbelen, Alexandre Piché, Bart Dhoedt, Aaron C. Courville, Alexandre Lacoste
Abstract
Controlling artificial agents from visual sensory data is an arduous task. Reinforcement learning (RL) algorithms can succeed but require large amounts of interactions between the agent and the environment. To alleviate the issue, unsupervised RL proposes to employ self-supervised interaction and learning, for adapting faster to future tasks. Yet, as shown in the Unsupervised RL Benchmark (URLB; Laskin et al. 2021), whether current unsupervised strategies can improve generalization capabilities is still unclear, especially in visual control settings. In this work, we study the URLB and propose a new method to solve it, using unsupervised model-based RL, for pre-training the agent, and a task-aware fine-tuning strategy combined with a new proposed hybrid planner, Dyna-MPC, to adapt the agent for downstream tasks. On URLB, our method obtains 93.59% overall normalized performance, surpassing previous baselines by a staggering margin. The approach is empirically evaluated through a large-scale empirical study, which we use to validate our design choices and analyze our models. We also show robust performance on the Real-Word RL benchmark, hinting at resiliency to environment perturbations during adaptation. Project website: https://masteringurlb.github.io/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7c415698-2ae1-4191-97b9-bc0b6ed5f3c6Cited by top-tier papers15
- METRA: Scalable Unsupervised RL with Metric-Aware AbstractionSeohong Park, Oleh Rybkin, Sergey LevineICLR 2024 · 83 citations
- Simple Hierarchical Planning with DiffusionChang Chen, Fei Deng, Kenji Kawaguchi, Caglar Gulcehre et al.ICLR 2024 · 79 citations
- Foundation Policies with Hilbert RepresentationsSeohong Park, Tobias Kreiman, Sergey LevineICML 2024 · 72 citations
- Emergent Dexterity Via Diverse Resets and Large-Scale Reinforcement LearningPatrick Yin, Tyler Westenbroek, Zhengyu Zhang, Ignacio Dagnino et al.ICLR 2026 · 15 citations
- PEAC: Unsupervised Pre-training for Cross-Embodiment Reinforcement LearningChengyang Ying, Zhongkai Hao, Xinning Zhou, Xuezhou Xu et al.NeurIPS 2024 · 14 citations
Builds on13
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel et al.ICML 2020 · 489 citations
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement LearningDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICLR 2022 · 457 citations
Related papers
- A Mixture Of Surprises for Unsupervised Reinforcement LearningAndrew Zhao, Matthieu Gaetan Lin, Yangguang Li, Yong-Jin Liu et al.NeurIPS 2022 · 15 citations
- EUCLID: Towards Efficient Unsupervised Reinforcement Learning with Multi-choice Dynamics ModelYifu Yuan, Jianye Hao, Fei Ni, Yao Mu et al.ICLR 2023 · 1 citation
- Unsupervised Reinforcement Learning with Contrastive Intrinsic ControlMichael Laskin, Hao Liu, Xue Bin Peng, Denis Yarats et al.NeurIPS 2022 · 62 citations
- Unsupervised Reinforcement Learning of Transferable Meta-Skills for Embodied NavigationJuncheng Li, Xin Wang, Siliang Tang, Haizhou Shi et al.CVPR 2020
- Self-Supervised Reinforcement Learning that Transfers using Random FeaturesBoyuan Chen, Chuning Zhu, Pulkit Agrawal, Kaiqing Zhang et al.NeurIPS 2023 · 16 citations
