Deep Surrogate Assisted Generation of Environments
Varun Bhatt, Bryon Tjanaka, Matthew C. Fontaine, Stefanos Nikolaidis
摘要
Recent progress in reinforcement learning (RL) has started producing generally capable agents that can solve a distribution of complex environments. These agents are typically tested on fixed, human-authored environments. On the other hand, quality diversity (QD) optimization has been proven to be an effective component of environment generation algorithms, which can generate collections of high-quality environments that are diverse in the resulting agent behaviors. However, these algorithms require potentially expensive simulations of agents on newly generated environments. We propose Deep Surrogate Assisted Generation of Environments (DSAGE), a sample-efficient QD environment generation algorithm that maintains a deep surrogate model for predicting agent behaviors in new environments. Results in two benchmark domains show that DSAGE significantly outperforms existing QD environment generation algorithms in discovering collections of environments that elicit diverse behaviors of a state-of-the-art RL agent and a planning agent. Our source code and videos are available at https://dsagepaper.github.io/ * Equal contribution Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Rainbow Teaming: Open-Ended Generation of Diverse Adversarial PromptsMikayel Samvelyan, Sharath Chandra Raparthy, Andrei Lupu, Eric Hambro 等NeurIPS 2024 · 被引用 231 次
- Human-Timescale Adaptation in an Open-Ended Task SpaceJakob Bauer, Kate Baumli, Feryal M. P. Behbahani, Avishkar Bhoopchand 等ICML 2023 · 被引用 155 次
- Quality-Diversity through AI FeedbackHerbie Bradley, Andrew Dai, Hannah Benita Teufel, Jenny Zhang 等ICLR 2024 · 被引用 43 次
- Robust Multi-Agent Coordination via Evolutionary Generation of Auxiliary Adversarial AttackersLei Yuan, Ziqian Zhang, Ke Xue, Hao Yin 等AAAI 2023 · 被引用 31 次
- Generating Behaviorally Diverse Policies with Latent Diffusion ModelsShashank Hegde, Sumeet Batra, K. R. Zentner, Gaurav S. SukhatmeNeurIPS 2023 · 被引用 27 次
它引用的顶会 Paper8
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 被引用 685 次
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen 等NeurIPS 2020 · 被引用 362 次
- Prioritized Level ReplayMinqi Jiang, Edward Grefenstette, Tim RocktäschelICML 2021 · 被引用 211 次
- Evolving Curricula with Regret-Based Environment DesignJack Parker-Holder, Minqi Jiang, Michael Dennis, Mikayel Samvelyan 等ICML 2022 · 被引用 175 次
- Enhanced POET: Open-ended Reinforcement Learning through Unbounded Invention of Learning Challenges and their SolutionsRui Wang, Joel Lehman, Aditya Rawal, Jiale Zhi 等ICML 2020 · 被引用 148 次
相关 Paper
- Enhancing Question Generation through Diversity-Seeking Reinforcement Learning with Bilevel Policy DecompositionTianyu Ren, Hui Wang, Karen RaffertyAAAI 2025 · 被引用 5 次
- AutoQD: Automatic Discovery of Diverse Behaviors with Quality-Diversity OptimizationSaeed Hedayatian, Stefanos NikolaidisICLR 2026 · 被引用 5 次
- Quality Diversity through Human Feedback: Towards Open-Ended Diversity-Driven OptimizationLi Ding, Jenny Zhang, Jeff Clune, Lee Spector 等ICML 2024 · 被引用 5 次
- Quality-Diversity Actor-Critic: Learning High-Performing and Diverse Behaviors via Value and Successor Features CriticsLuca Grillotti, Maxence Faldor, Borja G. León, Antoine CullyICML 2024 · 被引用 13 次
- Proximal Policy Gradient Arborescence for Quality Diversity Reinforcement LearningSumeet Batra, Bryon Tjanaka, Matthew Christopher Fontaine, Aleksei Petrenko 等ICLR 2024 · 被引用 26 次
