A Mixture Of Surprises for Unsupervised Reinforcement Learning
Andrew Zhao, Matthieu Gaetan Lin, Yangguang Li, Yong-Jin Liu, Gao Huang
Abstract
Unsupervised reinforcement learning aims at learning a generalist policy in a reward-free manner for fast adaptation to downstream tasks. Most of the existing methods propose to provide an intrinsic reward based on surprise. Maximizing or minimizing surprise drives the agent to either explore or gain control over its environment. However, both strategies rely on a strong assumption: the entropy of the environment's dynamics is either high or low. This assumption may not always hold in real-world scenarios, where the entropy of the environment's dynamics may be unknown. Hence, choosing between the two objectives is a dilemma. We propose a novel yet simple mixture of policies to address this concern, allowing us to optimize an objective that simultaneously maximizes and minimizes the surprise. Concretely, we train one mixture component whose objective is to maximize the surprise and another whose objective is to minimize the surprise. Hence, our method does not make assumptions about the entropy of the environment's dynamics. We call our method a (MOSS) for unsupervised reinforcement learning. Experimental results show that our simple method achieves state-of-the-art performance on the URLB benchmark, outperforming previous pure surprise maximization-based objectives. Our code is available at: https://github.com/LeapLabTHU/MOSS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f17dd5af-e159-4907-ad14-3a651d4226cbCited by top-tier papers10
- Absolute Zero: Reinforced Self-play Reasoning with Zero DataAndrew Zhao, Yiran Wu, Tong Wu, Quentin Xu et al.NeurIPS 2025 · 361 citations
- METRA: Scalable Unsupervised RL with Metric-Aware AbstractionSeohong Park, Oleh Rybkin, Sergey LevineICLR 2024 · 83 citations
- Controllability-Aware Unsupervised Skill DiscoverySeohong Park, Kimin Lee, Youngwoon Lee, Pieter AbbeelICML 2023 · 62 citations
- PEAC: Unsupervised Pre-training for Cross-Embodiment Reinforcement LearningChengyang Ying, Zhongkai Hao, Xinning Zhou, Xuezhou Xu et al.NeurIPS 2024 · 14 citations
- DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing ConstraintsAndrew Zhao, Quentin Xu, Matthieu Lin, Shenzhi Wang et al.AAAI 2025 · 11 citations
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar et al.ICLR 2020 · 475 citations
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement LearningDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICLR 2022 · 457 citations
- Reinforcement Learning with Prototypical RepresentationsDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICML 2021 · 262 citations
Related papers
- SMiRL: Surprise Minimizing Reinforcement Learning in Unstable EnvironmentsGlen Berseth, Daniel Geng, Coline Manon Devin, Nicholas Rhinehart et al.ICLR 2021 · 12 citations
- Discovering Mixture Skills for Unsupervised Reinforcement LearningNelson Ma, Junyu Xuan, Guangquan Zhang, Jie LuAAAI 2026
- Constrained Ensemble Exploration for Unsupervised Skill DiscoveryChenjia Bai, Rushuai Yang, Qiaosheng Zhang, Kang Xu et al.ICML 2024 · 9 citations
- Mastering the Unsupervised Reinforcement Learning Benchmark from PixelsSai Rajeswar, Pietro Mazzaglia, Tim Verbelen, Alexandre Piché et al.ICML 2023 · 30 citations
- Behavior Contrastive Learning for Unsupervised Skill DiscoveryRushuai Yang, Chenjia Bai, Hongyi Guo, Siyuan Li et al.ICML 2023 · 34 citations
