AutoQD: Automatic Discovery of Diverse Behaviors with Quality-Diversity Optimization
Saeed Hedayatian, Stefanos Nikolaidis
摘要
Quality-Diversity (QD) algorithms have shown remarkable success in discovering diverse, high-performing solutions, but rely heavily on hand-crafted behavioral descriptors that constrain exploration to predefined notions of diversity. Leveraging the equivalence between policies and occupancy measures, we present a theoretically grounded approach to automatically generate behavioral descriptors by embedding the occupancy measures of policies in Markov Decision Processes. Our method, AutoQD, leverages random Fourier features to approximate the Maximum Mean Discrepancy (MMD) between policy occupancy measures, creating embeddings whose distances reflect meaningful behavioral differences. A low-dimensional projection of these embeddings that captures the most behaviorally significant dimensions can then be used as behavioral descriptors for CMA-MAE, a state of the art blackbox QD method, to discover diverse policies. We prove that our embeddings converge to true MMD distances between occupancy measures as the number of sampled trajectories and embedding dimensions increase. Through experiments in multiple continuous control tasks we demonstrate AutoQD's ability in discovering diverse policies without predefined behavioral descriptors, presenting a well-motivated alternative to prior methods in unsupervised Reinforcement Learning and QD optimization. Our approach opens new possibilities for open-ended learning and automated behavior discovery in sequential decision making settings without requiring domain-specific knowledge.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper13
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar 等ICLR 2020 · 被引用 475 次
- Effective Diversity in Population Based Reinforcement LearningJack Parker-Holder, Aldo Pacchiano, Krzysztof Marcin Choromanski, Stephen J. RobertsNeurIPS 2020 · 被引用 195 次
- Differentiable Quality DiversityMatthew C. Fontaine, Stefanos NikolaidisNeurIPS 2021 · 被引用 116 次
- One Solution is Not All You Need: Few-Shot Extrapolation via Structured MaxEnt RLSaurabh Kumar, Aviral Kumar, Sergey Levine, Chelsea FinnNeurIPS 2020 · 被引用 109 次
- Illuminating Mario Scenes in the Latent Space of a Generative Adversarial NetworkMatthew C. Fontaine, Ruilin Liu, Ahmed Khalifa, Jignesh Modi 等AAAI 2021 · 被引用 98 次
相关 Paper
- Discount Model Search for Quality Diversity Optimization in High-Dimensional Measure SpacesBryon Tjanaka, Henry Chen, Matthew Christopher Fontaine, Stefanos NikolaidisICLR 2026 · 被引用 1 次
- Discovering Policies with DOMiNO: Diversity Optimization Maintaining Near OptimalityTom Zahavy, Yannick Schroecker, Feryal M. P. Behbahani, Kate Baumli 等ICLR 2023 · 被引用 2 次
- Deep Surrogate Assisted Generation of EnvironmentsVarun Bhatt, Bryon Tjanaka, Matthew C. Fontaine, Stefanos NikolaidisNeurIPS 2022 · 被引用 54 次
- Quality-Similar Diversity via Population Based Reinforcement LearningShuang Wu, Jian Yao, Haobo Fu, Ye Tian 等ICLR 2023
- Towards Unifying Behavioral and Response Diversity for Open-ended Learning in Zero-sum GamesXiangyu Liu, Hangtian Jia, Ying Wen, Yujing Hu 等NeurIPS 2021 · 被引用 67 次
