Negatively Correlated Ensemble Reinforcement Learning for Online Diverse Game Level Generation
Ziqi Wang, Chengpeng Hu, Jialin Liu, Xin Yao
Abstract
Deep reinforcement learning has recently been successfully applied to online procedural content generation in which a policy determines promising game-level segments. However, existing methods can hardly discover diverse level patterns, while the lack of diversity makes the gameplay boring. This paper proposes an ensemble reinforcement learning approach that uses multiple negatively correlated sub-policies to generate different alternative level segments, and stochastically selects one of them following a dynamic selector policy. A novel policy regularisation technique is integrated into the approach to diversify the generated alternatives. In addition, we develop theorems to provide general methodologies for optimising policy regularisation in a Markov decision process. The proposed approach is compared with several state-of-the-art policy ensemble methods and classic methods on a well-known level generation benchmark, with two different reward functions expressing game-design goals from different perspectives. Results show that our approach boosts level diversity notably with competitive performance in terms of the reward. Furthermore, by varying the regularisation coefficient values, the trained generators form a well-spread Pareto front, allowing explicit trade-offs between diversity and rewards of generated levels.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1d554b2f-4df9-4224-a5fa-8439bdbeba53Cited by top-tier papers2
- Learning Intractable Multimodal Policies with Reparameterization and Diversity RegularizationZiqi Wang, Jiashun Liu, Ling PanNeurIPS 2025 · 3 citations
- Reinforcement Learning with Adaptive Reward Modeling for Expensive-to-Evaluate SystemsHongyuan Su, Yu Zheng, Yuan Yuan, Yuming Lin et al.ICML 2025
Builds on7
- SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement LearningKimin Lee, Michael Laskin, Aravind Srinivas, Pieter AbbeelICML 2021 · 239 citations
- Effective Diversity in Population Based Reinforcement LearningJack Parker-Holder, Aldo Pacchiano, Krzysztof Marcin Choromanski, Stephen J. RobertsNeurIPS 2020 · 195 citations
- Trajectory Diversity for Zero-Shot CoordinationAndrei Lupu, Brandon Cui, Hengyuan Hu, Jakob N. FoersterICML 2021 · 157 citations
- Illuminating Mario Scenes in the Latent Space of a Generative Adversarial NetworkMatthew C. Fontaine, Ruilin Liu, Ahmed Khalifa, Jignesh Modi et al.AAAI 2021 · 98 citations
- Maximizing Ensemble Diversity in Deep Reinforcement LearningHassam Sheikh, Mariano Phielipp, Ladislau BölöniICLR 2022 · 10 citations
Related papers
- Ensemble-based Deep Reinforcement Learning for Vehicle Routing Problems under Distribution ShiftYuan Jiang, Zhiguang Cao, Yaoxin Wu, Wen Song et al.NeurIPS 2023 · 43 citations
- Automatic Data Augmentation for Generalization in Reinforcement LearningRoberta Raileanu, Maxwell Goldstein, Denis Yarats, Ilya Kostrikov et al.NeurIPS 2021 · 143 citations
- Entropy-regularized Diffusion Policy with Q-Ensembles for Offline Reinforcement LearningRuoqi Zhang, Ziwei Luo, Jens Sjölund, Thomas B. Schön et al.NeurIPS 2024 · 43 citations
- Rethinking Policy Diversity in Ensemble Policy Gradient in Large-Scale Reinforcement LearningNaoki Shitanda, Motoki Omura, Tatsuya Harada, Takayuki OsaICLR 2026
- DiverseGRPO: Mitigating Mode Collapse in Image Generation via Diversity-Aware GRPOHenglin Liu, Huijuan Huang, Jing Wang, Chang Liu et al.CVPR 2026 · 17 citations
