Learning General World Models in a Handful of Reward-Free Deployments
Yingchen Xu, Jack Parker-Holder, Aldo Pacchiano, Philip J. Ball, Oleh Rybkin, Stephen Roberts, Tim Rocktäschel, Edward Grefenstette
Abstract
Building generally capable agents is a grand challenge for deep reinforcement learning (RL). To approach this challenge practically, we outline two key desiderata: 1) to facilitate generalization, exploration should be task agnostic; 2) to facilitate scalability, exploration policies should collect large quantities of data without costly centralized retraining. Combining these two properties, we introduce the reward-free deployment efficiency setting, a new paradigm for RL research. We then present CASCADE, a novel approach for self-supervised exploration in this new setting. CASCADE seeks to learn a world model by collecting data with a population of agents, using an information theoretic objective inspired by Bayesian Active Learning. CASCADE achieves this by specifically maximizing the diversity of trajectories sampled by the population through a novel cascading objective. We provide theoretical intuition for CASCADE which we show in a tabular setting improves upon naïve approaches that do not account for population diversity. We then demonstrate that CASCADE collects diverse task-agnostic datasets and learns agents that generalize zero-shot to novel, unseen downstream tasks on Atari, MiniGrid, Crafter and the DM Control Suite. Code and videos are available at https://ycxuyingchen.github.io/cascade/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Provable Reward-Agnostic Preference-Based Reinforcement LearningWenhao Zhan, Masatoshi Uehara, Wen Sun, Jason D. LeeICLR 2024 · 16 citations
- Reward-Free Curricula for Training Robust World ModelsMarc Rigter, Minqi Jiang, Ingmar PosnerICLR 2024 · 13 citations
- Distributional Successor Features Enable Zero-Shot Policy OptimizationChuning Zhu, Xinqi Wang, Tyler Han, Simon S. Du et al.NeurIPS 2024 · 11 citations
- Efficient Reinforcement Learning by Guiding World Models with Non-Curated DataYi Zhao, Aidan Scannell, Wenshuai Zhao, Yuxin Hou et al.ICLR 2026 · 2 citations
- DreamSAC: Learning Hamiltonian World Models via Symmetry ExplorationJinzhou Tang, Fan Feng, Minghao Fu, Wenjun Lin et al.CVPR 2026 · 1 citation
Builds on46
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
Related papers
- Towards Principled Unsupervised Multi-Agent Reinforcement LearningRiccardo Zamboni, Mirco Mutti, Marcello RestelliNeurIPS 2025 · 5 citations
- Task-agnostic Exploration in Reinforcement LearningXuezhou Zhang, Yuzhe Ma, Adish SinglaNeurIPS 2020 · 56 citations
- Meta-DT: Offline Meta-RL as Conditional Sequence Modeling with World Model DisentanglementZhi Wang, Li Zhang, Wenhao Wu, Yuanheng Zhu et al.NeurIPS 2024 · 31 citations
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel et al.ICML 2020 · 489 citations
- CUDC: A Curiosity-Driven Unsupervised Data Collection Method with Adaptive Temporal Distances for Offline Reinforcement LearningChenyu Sun, Hangwei Qian, Chunyan MiaoAAAI 2024 · 1 citation
