DGPO: Discovering Multiple Strategies with Diversity-Guided Policy Optimization
Wentse Chen, Shiyu Huang, Yuan Chiang, Tim Pearce, Wei-Wei Tu, Ting Chen, Jun Zhu
摘要
Most reinforcement learning algorithms seek a single optimal strategy that solves a given task. However, it can often be valuable to learn a diverse set of solutions, for instance, to make an agent's interaction with users more engaging, or improve the robustness of a policy to an unexpected perturbance. We propose Diversity-Guided Policy Optimization (DGPO), an on-policy algorithm that discovers multiple strategies for solving a given task. Unlike prior work, it achieves this with a shared policy network trained over a single run. Specifically, we design an intrinsic reward based on an information-theoretic diversity objective. Our final objective alternately constraints on the diversity of the strategies and on the extrinsic reward. We solve the constrained optimization problem by casting it as a probabilistic inference task and use policy iteration to maximize the derived lower bound. Experimental results show that our method efficiently discovers diverse strategies in a wide variety of reinforcement learning tasks. Compared to baseline methods, DGPO achieves comparable rewards, while discovering more diverse strategies, and often with better sample efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper10
- Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable ModelAlex X. Lee, Anusha Nagabandi, Pieter Abbeel, Sergey LevineNeurIPS 2020 · 被引用 437 次
- Effective Diversity in Population Based Reinforcement LearningJack Parker-Holder, Aldo Pacchiano, Krzysztof Marcin Choromanski, Stephen J. RobertsNeurIPS 2020 · 被引用 195 次
- One Solution is Not All You Need: Few-Shot Extrapolation via Structured MaxEnt RLSaurabh Kumar, Aviral Kumar, Sergey Levine, Chelsea FinnNeurIPS 2020 · 被引用 109 次
- Learning to Coordinate Manipulation Skills via Skill Behavior DiversificationYoungwoon Lee, Jingyun Yang, Joseph J. LimICLR 2020 · 被引用 98 次
- Discovering Diverse Multi-Agent Strategic Behavior via Reward RandomizationZhenggang Tang, Chao Yu, Boyuan Chen, Huazhe Xu 等ICLR 2021 · 被引用 63 次
相关 Paper
- Continuously Discovering Novel Strategies via Reward-Switching Policy OptimizationZihan Zhou, Wei Fu, Bingliang Zhang, Yi WuICLR 2022 · 被引用 34 次
- Discovering Policies with DOMiNO: Diversity Optimization Maintaining Near OptimalityTom Zahavy, Yannick Schroecker, Feryal M. P. Behbahani, Kate Baumli 等ICLR 2023 · 被引用 2 次
- Iteratively Learn Diverse Strategies with State Distance InformationWei Fu, Weihua Du, Jingwei Li, Sunli Chen 等NeurIPS 2023 · 被引用 5 次
- Controlling Behavioral Diversity in Multi-Agent Reinforcement LearningMatteo Bettini, Ryan Kortvelesy, Amanda ProrokICML 2024 · 被引用 11 次
- Generalized Behavior Learning from Diverse DemonstrationsVarshith Sreeramdass, Rohan R. Paleja, Letian Chen, Sanne van Waveren 等ICLR 2025
