DRED: Zero-Shot Transfer in Reinforcement Learning via Data-Regularised Environment Design
Samuel Garcin, James Doran, Shangmin Guo, Christopher G. Lucas, Stefano V. Albrecht
Abstract
Autonomous agents trained using deep reinforcement learning (RL) often lack the ability to successfully generalise to new environments, even when these environments share characteristics with the ones they have encountered during training. In this work, we investigate how the sampling of individual environment instances, or levels, affects the zero-shot generalisation (ZSG) ability of RL agents. We discover that, for deep actor-critic architectures sharing their base layers, prioritising levels according to their value loss minimises the mutual information between the agent's internal representation and the set of training levels in the generated training data. This provides a novel theoretical justification for the regularisation achieved by certain adaptive sampling strategies. We then turn our attention to unsupervised environment design (UED) methods, which assume control over level generation. We find that existing UED methods can significantly shift the training distribution, which translates to low ZSG performance. To prevent both overfitting and distributional shift, we introduce data-regularised environment design (DRED). DRED generates levels using a generative model trained to approximate the ground truth distribution of an initial set of level parameters. Through its grounding, DRED achieves significant improvements in ZSG over adaptive level sampling strategies and UED methods. Our code and experimental data are available at https://github.com/uoe-agents/dred.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68d823fa-1776-4843-a600-2b6c76b87cbeCited by top-tier papers9
- No Regrets: Investigating and Improving Regret Approximations for Curriculum DiscoveryAlexander Rutherford, Michael Beukman, Timon Willi, Bruno Lacerda et al.NeurIPS 2024 · 38 citations
- Refining Minimax Regret for Unsupervised Environment DesignMichael Beukman, Samuel Coward, Michael T. Matthews, Mattie Fellows et al.ICML 2024 · 15 citations
- Adversarial Environment Design via Regret-Guided Diffusion ModelsHojun Chung, Junseo Lee, Minsoo Kim, Dohyeong Kim et al.NeurIPS 2024 · 11 citations
- Skill-aware Mutual Information Optimisation for Zero-shot Generalisation in Reinforcement LearningXuehui Yu, Mhairi Dunion, Xin Li, Stefano V. AlbrechtNeurIPS 2024 · 6 citations
- Beyond Pixel Histories: World Models with Persistent 3D StateSamuel Garcin, Tom Walker, Steven McDonagh, Tim Pearce et al.ICML 2026 · 6 citations
Builds on15
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 911 citations
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 685 citations
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen et al.NeurIPS 2020 · 362 citations
- Prioritized Level ReplayMinqi Jiang, Edward Grefenstette, Tim RocktäschelICML 2021 · 211 citations
Related papers
- Replay-Guided Adversarial Environment DesignMinqi Jiang, Michael Dennis, Jack Parker-Holder, Jakob N. Foerster et al.NeurIPS 2021 · 148 citations
- Improving Regret Approximation for Unsupervised Dynamic Environment GenerationHarry Mead, Bruno Lacerda, Jakob N. Foerster, Nick HawesNeurIPS 2025 · 1 citation
- PACE: Parameter Change for Unsupervised Environment DesignFang YUAN, Junjie Zeng, Qinglun Li, Long Qin et al.ICML 2026
- CLUTR: Curriculum Learning via Unsupervised Task Representation LearningAbdus Salam Azad, Izzeddin Gur, Jasper Emhoff, Nathaniel Alexis et al.ICML 2023 · 20 citations
- Discovering General Reinforcement Learning Algorithms with Adversarial Environment DesignMatthew Thomas Jackson, Minqi Jiang, Jack Parker-Holder, Risto Vuorio et al.NeurIPS 2023 · 23 citations
