Quality-Diversity Actor-Critic: Learning High-Performing and Diverse Behaviors via Value and Successor Features Critics
Luca Grillotti, Maxence Faldor, Borja G. León, Antoine Cully
摘要
A key aspect of intelligence is the ability to demonstrate a broad spectrum of behaviors for adapting to unexpected situations. Over the past decade, advancements in deep reinforcement learning have led to groundbreaking achievements to solve complex continuous control tasks. However, most approaches return only one solution specialized for a specific problem. We introduce Quality-Diversity Actor-Critic (QDAC), an off-policy actor-critic deep reinforcement learning algorithm that leverages a value function critic and a successor features critic to learn high-performing and diverse behaviors. In this framework, the actor optimizes an objective that seamlessly unifies both critics using constrained optimization to (1) maximize return, while (2) executing diverse skills. Compared with other Quality-Diversity methods, QDAC achieves significantly higher performance and more diverse behaviors on six challenging continuous control locomotion tasks. We also demonstrate that we can harness the learned skills to adapt better than other baselines to five perturbed environments. Finally, qualitative analyses showcase a range of remarkable behaviors: adaptive-intelligent-robotics.github.io/QDAC.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Diversity-Incentivized Exploration for Versatile ReasoningZican Hu, Shilin Zhang, Yafu Li, Jianhao Yan 等ICLR 2026 · 被引用 32 次
- Zero-Shot Adaptation of Behavioral Foundation Models to Unseen DynamicsMaksim Bobrin, Ilya Zisman, Alexander Nikulin, Vladislav Kurenkov 等ICLR 2026 · 被引用 9 次
- Walk Wisely on Graph: Knowledge Graph Reasoning with Dual Agents via Efficient Guidance-ExplorationZijian Wang, Bin Wang, Haifeng Jing, Huayu Li 等AAAI 2025 · 被引用 6 次
- Is On-Policy Data always the Best Choice for Direct Preference Optimization-Based LM Alignment?Zetian Sun, Dongfang Li, Xuhui Chen, Baotian Hu 等ICLR 2026 · 被引用 1 次
- Discount Model Search for Quality Diversity Optimization in High-Dimensional Measure SpacesBryon Tjanaka, Henry Chen, Matthew Christopher Fontaine, Stefanos NikolaidisICLR 2026 · 被引用 1 次
它引用的顶会 Paper11
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar 等ICLR 2020 · 被引用 475 次
- Behavior From the Void: Unsupervised Active Pre-TrainingHao Liu, Pieter AbbeelNeurIPS 2021 · 被引用 258 次
- One Solution is Not All You Need: Few-Shot Extrapolation via Structured MaxEnt RLSaurabh Kumar, Aviral Kumar, Sergey Levine, Chelsea FinnNeurIPS 2020 · 被引用 109 次
相关 Paper
- Diversifying Policy Behaviors with Extrinsic Behavioral CuriosityZhenglin Wan, Xingrui Yu, David Mark Bossens, Yueming Lyu 等ICML 2025
- Proximal Policy Gradient Arborescence for Quality Diversity Reinforcement LearningSumeet Batra, Bryon Tjanaka, Matthew Christopher Fontaine, Aleksei Petrenko 等ICLR 2024 · 被引用 26 次
- Neuroevolution is a Competitive Alternative to Reinforcement Learning for Skill DiscoveryFélix Chalumeau, Raphaël Boige, Bryan Lim, Valentin Macé 等ICLR 2023 · 被引用 6 次
- Discovering Policies with DOMiNO: Diversity Optimization Maintaining Near OptimalityTom Zahavy, Yannick Schroecker, Feryal M. P. Behbahani, Kate Baumli 等ICLR 2023 · 被引用 2 次
- Decoupling Value and Policy for Generalization in Reinforcement LearningRoberta Raileanu, Rob FergusICML 2021 · 被引用 116 次
