PACE: Parameter Change for Unsupervised Environment Design
Fang YUAN, Junjie Zeng, Qinglun Li, Long Qin, Quanjun Yin, Siqi Shen, Yuxiang Xie, Junqiang Yang
Abstract
Unsupervised Environment Design (UED) offers a promising paradigm for improving reinforcement learning generalization by adaptively shaping training environments, but it requires reliable environment evaluation to remain effective. However, existing UED methods evaluate environments using indirect proxy signals such as regret, value-based errors, or Monte Carlo, which suffer from bias, high variance, or substantial computational overhead. To address these limitations, we propose Parameter Change Environment Design (PACE), a general framework for adaptive level selection in UED. PACE evaluates a level by performing a provisional policy update on it and scoring it with the squared norm of the induced parameter change, which directly reflects realized learning progress. This score then guides level selection: levels enter a staleness-aware buffer based on their score, and are replayed via rank-based prioritization, inducing a curriculum that adapts to the agent’s evolving capability. By grounding environment evaluation in intrinsic optimization progress, PACE provides a low-variance evaluation signal and avoids the need for additional environment rollouts. Experiments on MiniGrid and Craftax demonstrate that PACE consistently outperforms established UED baselines in zero-shot generalization across diverse out-of-distribution evaluation protocols.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e1e76dab-85ac-4eb5-96c7-ac1214dd9c53Builds on12
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen et al.NeurIPS 2020 · 362 citations
- The NetHack Learning EnvironmentHeinrich Küttler, Nantas Nardelli, Alexander H. Miller, Roberta Raileanu et al.NeurIPS 2020 · 251 citations
- Prioritized Level ReplayMinqi Jiang, Edward Grefenstette, Tim RocktäschelICML 2021 · 211 citations
- Benchmarking the Spectrum of Agent CapabilitiesDanijar HafnerICLR 2022 · 193 citations
Related papers
- Replay-Guided Adversarial Environment DesignMinqi Jiang, Michael Dennis, Jack Parker-Holder, Jakob N. Foerster et al.NeurIPS 2021 · 148 citations
- CLUTR: Curriculum Learning via Unsupervised Task Representation LearningAbdus Salam Azad, Izzeddin Gur, Jasper Emhoff, Nathaniel Alexis et al.ICML 2023 · 20 citations
- TRACED: Transition-aware Regret Approximation with Co-learnability for Environment DesignGeonwoo Cho, Jaegyun Im, Jihwan Lee, Hojun Yi et al.ICLR 2026 · 1 citation
- No Regrets: Investigating and Improving Regret Approximations for Curriculum DiscoveryAlexander Rutherford, Michael Beukman, Timon Willi, Bruno Lacerda et al.NeurIPS 2024 · 38 citations
- Improving Regret Approximation for Unsupervised Dynamic Environment GenerationHarry Mead, Bruno Lacerda, Jakob N. Foerster, Nick HawesNeurIPS 2025 · 1 citation
