Proximal Policy Gradient Arborescence for Quality Diversity Reinforcement Learning
Sumeet Batra, Bryon Tjanaka, Matthew Christopher Fontaine, Aleksei Petrenko, Stefanos Nikolaidis, Gaurav S. Sukhatme
Abstract
Training generally capable agents that thoroughly explore their environment and learn new and diverse skills is a long-term goal of robot learning. Quality Diversity Reinforcement Learning (QD-RL) is an emerging research area that blends the best aspects of both fields -- Quality Diversity (QD) provides a principled form of exploration and produces collections of behaviorally diverse agents, while Reinforcement Learning (RL) provides a powerful performance improvement operator enabling generalization across tasks and dynamic environments. Existing QD-RL approaches have been constrained to sample efficient, deterministic off-policy RL algorithms and/or evolution strategies, and struggle with highly stochastic environments. In this work, we, for the first time, adapt on-policy RL, specifically Proximal Policy Optimization (PPO), to the Differentiable Quality Diversity (DQD) framework and propose additional improvements over prior work that enable efficient optimization and discovery of novel skills on challenging locomotion tasks. Our new algorithm, Proximal Policy Gradient Arborescence (PPGA), achieves state-of-the-art results, including a 4x improvement in best reward over baselines on the challenging humanoid domain.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5cd18939-668d-4cb6-93d2-a3b0cfa1736eCited by top-tier papers9
- Diversity-Aware Policy Optimization for Large Language Model ReasoningJian Yao, Ran Cheng, Xingyu Wu, Jibin Wu et al.NeurIPS 2025 · 44 citations
- Sample-Efficient Quality-Diversity by Cooperative CoevolutionKe Xue, Ren-Jian Wang, Pengyi Li, Dong Li et al.ICLR 2024 · 17 citations
- Quality-Diversity Actor-Critic: Learning High-Performing and Diverse Behaviors via Value and Successor Features CriticsLuca Grillotti, Maxence Faldor, Borja G. León, Antoine CullyICML 2024 · 13 citations
- AutoQD: Automatic Discovery of Diverse Behaviors with Quality-Diversity OptimizationSaeed Hedayatian, Stefanos NikolaidisICLR 2026 · 5 citations
- POI Recommendation via Multi-Objective Adversarial Imitation LearningZhenglin Wan, Anjun Gao, Xingrui Yu, Pingfu Chao et al.AAAI 2025 · 1 citation
Builds on2
Related papers
- Diversifying Policy Behaviors with Extrinsic Behavioral CuriosityZhenglin Wan, Xingrui Yu, David Mark Bossens, Yueming Lyu et al.ICML 2025
- Polychromic Objectives for Reinforcement LearningJubayer Ibn Hamid, Ifdita Hasan Orney, Ellen Xu, Chelsea Finn et al.ICLR 2026 · 9 citations
- Guided Exploration with Proximal Policy Optimization using a Single DemonstrationGabriele Libardi, Gianni De Fabritiis, Sebastian DittertICML 2021 · 32 citations
- Neuroevolution is a Competitive Alternative to Reinforcement Learning for Skill DiscoveryFélix Chalumeau, Raphaël Boige, Bryan Lim, Valentin Macé et al.ICLR 2023 · 6 citations
- Sub-policy Adaptation for Hierarchical Reinforcement LearningAlexander C. Li, Carlos Florensa, Ignasi Clavera, Pieter AbbeelICLR 2020 · 85 citations
