PACE: Parameter Change for Unsupervised Environment Design
Fang YUAN, Junjie Zeng, Qinglun Li, Long Qin, Quanjun Yin, Siqi Shen, Yuxiang Xie, Junqiang Yang
摘要
Unsupervised Environment Design (UED) offers a promising paradigm for improving reinforcement learning generalization by adaptively shaping training environments, but it requires reliable environment evaluation to remain effective. However, existing UED methods evaluate environments using indirect proxy signals such as regret, value-based errors, or Monte Carlo, which suffer from bias, high variance, or substantial computational overhead. To address these limitations, we propose Parameter Change Environment Design (PACE), a general framework for adaptive level selection in UED. PACE evaluates a level by performing a provisional policy update on it and scoring it with the squared norm of the induced parameter change, which directly reflects realized learning progress. This score then guides level selection: levels enter a staleness-aware buffer based on their score, and are replayed via rank-based prioritization, inducing a curriculum that adapts to the agent’s evolving capability. By grounding environment evaluation in intrinsic optimization progress, PACE provides a low-variance evaluation signal and avoids the need for additional environment rollouts. Experiments on MiniGrid and Craftax demonstrate that PACE consistently outperforms established UED baselines in zero-shot generalization across diverse out-of-distribution evaluation protocols.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen 等NeurIPS 2020 · 被引用 362 次
- The NetHack Learning EnvironmentHeinrich Küttler, Nantas Nardelli, Alexander H. Miller, Roberta Raileanu 等NeurIPS 2020 · 被引用 251 次
- Prioritized Level ReplayMinqi Jiang, Edward Grefenstette, Tim RocktäschelICML 2021 · 被引用 211 次
- Benchmarking the Spectrum of Agent CapabilitiesDanijar HafnerICLR 2022 · 被引用 193 次
相关 Paper
- Replay-Guided Adversarial Environment DesignMinqi Jiang, Michael Dennis, Jack Parker-Holder, Jakob N. Foerster 等NeurIPS 2021 · 被引用 148 次
- CLUTR: Curriculum Learning via Unsupervised Task Representation LearningAbdus Salam Azad, Izzeddin Gur, Jasper Emhoff, Nathaniel Alexis 等ICML 2023 · 被引用 20 次
- TRACED: Transition-aware Regret Approximation with Co-learnability for Environment DesignGeonwoo Cho, Jaegyun Im, Jihwan Lee, Hojun Yi 等ICLR 2026 · 被引用 1 次
- No Regrets: Investigating and Improving Regret Approximations for Curriculum DiscoveryAlexander Rutherford, Michael Beukman, Timon Willi, Bruno Lacerda 等NeurIPS 2024 · 被引用 38 次
- Improving Regret Approximation for Unsupervised Dynamic Environment GenerationHarry Mead, Bruno Lacerda, Jakob N. Foerster, Nick HawesNeurIPS 2025 · 被引用 1 次
