Refining Minimax Regret for Unsupervised Environment Design
Michael Beukman, Samuel Coward, Michael T. Matthews, Mattie Fellows, Minqi Jiang, Michael D. Dennis, Jakob Nicolaus Foerster
摘要
In unsupervised environment design, reinforcement learning agents are trained on environment configurations (levels) generated by an adversary that maximises some objective. Regret is a commonly used objective that theoretically results in a minimax regret (MMR) policy with desirable robustness guarantees; in particular, the agent's maximum regret is bounded. However, once the agent reaches this regret bound on all levels, the adversary will only sample levels where regret cannot be further reduced. Although there may be possible performance improvements to be made outside of these regret-maximising levels, learning stagnates. In this work, we introduce Bayesian level-perfect MMR (BLP), a refinement of the minimax regret objective that overcomes this limitation. We formally show that solving for this objective results in a subset of MMR policies, and that BLP policies act consistently with a Perfect Bayesian policy over all levels. We further introduce an algorithm, ReMiDi, that results in a BLP policy at convergence. We empirically demonstrate that training on levels from a minimax regret adversary causes learning to prematurely stagnate, but that ReMiDi continues learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- No Regrets: Investigating and Improving Regret Approximations for Curriculum DiscoveryAlexander Rutherford, Michael Beukman, Timon Willi, Bruno Lacerda 等NeurIPS 2024 · 被引用 38 次
- Improving Environment Novelty Quantification for Effective Unsupervised Environment DesignJayden Teoh, Wenjun Li, Pradeep VarakanthamNeurIPS 2024 · 被引用 6 次
- Leveraging Separated World Model for Exploration in Visually Distracted EnvironmentsKaichen Huang, Shenghua Wan, Minghao Shao, Hai-Hang Sun 等NeurIPS 2024 · 被引用 5 次
- Unsupervised Partner Design Enables Robust Ad-hoc TeamworkConstantin Ruhdorfer, Matteo Bortoletto, Victor Oei, Anna Penzkofer 等ICML 2026 · 被引用 3 次
- Reward-Guided Prompt Evolving in Reinforcement Learning for LLMsZiyu Ye, Rishabh Agarwal, Tianqi Liu, Rishabh Joshi 等ICML 2025
它引用的顶会 Paper18
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen 等NeurIPS 2020 · 被引用 362 次
- "Other-Play" for Zero-Shot CoordinationHengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob N. FoersterICML 2020 · 被引用 271 次
- Prioritized Level ReplayMinqi Jiang, Edward Grefenstette, Tim RocktäschelICML 2021 · 被引用 211 次
- Evolving Curricula with Regret-Based Environment DesignJack Parker-Holder, Minqi Jiang, Michael Dennis, Mikayel Samvelyan 等ICML 2022 · 被引用 175 次
- Human-Timescale Adaptation in an Open-Ended Task SpaceJakob Bauer, Kate Baumli, Feryal M. P. Behbahani, Avishkar Bhoopchand 等ICML 2023 · 被引用 155 次
相关 Paper
- EUBRL: Epistemic Uncertainty Directed Bayesian Reinforcement LearningJianfei Ma, Wee Sun LeeICLR 2026 · 被引用 2 次
- R2-B2: Recursive Reasoning-Based Bayesian Optimization for No-Regret Learning in GamesZhongxiang Dai, Yizhou Chen, Bryan Kian Hsiang Low, Patrick Jaillet 等ICML 2020 · 被引用 28 次
- Replay-Guided Adversarial Environment DesignMinqi Jiang, Michael Dennis, Jack Parker-Holder, Jakob N. Foerster 等NeurIPS 2021 · 被引用 148 次
- Bayesian Robust Cooperative Multi-Agent Reinforcement Learning Against Unknown AdversariesKiarash Kazari, György DánICLR 2026
- Grounding Aleatoric Uncertainty for Unsupervised Environment DesignMinqi Jiang, Michael Dennis, Jack Parker-Holder, Andrei Lupu 等NeurIPS 2022 · 被引用 18 次
