MAESTRO: Open-Ended Environment Design for Multi-Agent Reinforcement Learning
Mikayel Samvelyan, Akbir Khan, Michael Dennis, Minqi Jiang, Jack Parker-Holder, Jakob Nicolaus Foerster, Roberta Raileanu, Tim Rocktäschel
Abstract
Open-ended learning methods that automatically generate a curriculum of increasingly challenging tasks serve as a promising avenue toward generally capable reinforcement learning agents. Existing methods adapt curricula independently over either environment parameters (in single-agent settings) or co-player policies (in multi-agent settings). However, the strengths and weaknesses of co-players can manifest themselves differently depending on environmental features. It is thus crucial to consider the dependency between the environment and co-player when shaping a curriculum in multi-agent domains. In this work, we use this insight and extend Unsupervised Environment Design (UED) to multi-agent environments. We then introduce Multi-Agent Environment Design Strategist for Open-Ended Learning (MAESTRO), the first multi-agent UED approach for two-player zero-sum settings. MAESTRO efficiently produces adversarial, joint curricula over both environments and co-players and attains minimax-regret guarantees at Nash equilibrium. Our experiments show that MAESTRO outperforms a number of strong baselines on competitive two-player games, spanning discrete and continuous control settings. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4cedae58-a03f-44eb-b063-d6d0c2d8b7a8Cited by top-tier papers15
- Rainbow Teaming: Open-Ended Generation of Diverse Adversarial PromptsMikayel Samvelyan, Sharath Chandra Raparthy, Andrei Lupu, Eric Hambro et al.NeurIPS 2024 · 231 citations
- Human-Timescale Adaptation in an Open-Ended Task SpaceJakob Bauer, Kate Baumli, Feryal M. P. Behbahani, Avishkar Bhoopchand et al.ICML 2023 · 155 citations
- No Regrets: Investigating and Improving Regret Approximations for Curriculum DiscoveryAlexander Rutherford, Michael Beukman, Timon Willi, Bruno Lacerda et al.NeurIPS 2024 · 38 citations
- Discovering General Reinforcement Learning Algorithms with Adversarial Environment DesignMatthew Thomas Jackson, Minqi Jiang, Jack Parker-Holder, Risto Vuorio et al.NeurIPS 2023 · 23 citations
- Refining Minimax Regret for Unsupervised Environment DesignMichael Beukman, Samuel Coward, Michael T. Matthews, Mattie Fellows et al.ICML 2024 · 15 citations
Builds on11
- Emergent Tool Use From Multi-Agent AutocurriculaBowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu et al.ICLR 2020 · 751 citations
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen et al.NeurIPS 2020 · 362 citations
- Prioritized Level ReplayMinqi Jiang, Edward Grefenstette, Tim RocktäschelICML 2021 · 211 citations
- Evolving Curricula with Regret-Based Environment DesignJack Parker-Holder, Minqi Jiang, Michael Dennis, Mikayel Samvelyan et al.ICML 2022 · 175 citations
- Enhanced POET: Open-ended Reinforcement Learning through Unbounded Invention of Learning Challenges and their SolutionsRui Wang, Joel Lehman, Aditya Rawal, Jiale Zhi et al.ICML 2020 · 148 citations
Related papers
- Improving Regret Approximation for Unsupervised Dynamic Environment GenerationHarry Mead, Bruno Lacerda, Jakob N. Foerster, Nick HawesNeurIPS 2025 · 1 citation
- Replay-Guided Adversarial Environment DesignMinqi Jiang, Michael Dennis, Jack Parker-Holder, Jakob N. Foerster et al.NeurIPS 2021 · 148 citations
- Beyond Fixed Tasks: Unsupervised Environment Design for Task-Level PairsDaniel Furelos-Blanco, Charles Pert, Frederik Kelbel, Alex F. Spies et al.AAAI 2026
- It Takes Four to Tango: Multiagent Self Play for Automatic Curriculum GenerationYuqing Du, Pieter Abbeel, Aditya GroverICLR 2022 · 20 citations
- Adversarial Environment Design via Regret-Guided Diffusion ModelsHojun Chung, Junseo Lee, Minsoo Kim, Dohyeong Kim et al.NeurIPS 2024 · 11 citations
