Iteratively Learn Diverse Strategies with State Distance Information
Wei Fu, Weihua Du, Jingwei Li, Sunli Chen, Jingzhao Zhang, Yi Wu
摘要
In complex reinforcement learning (RL) problems, policies with similar rewards may have substantially different behaviors. It remains a fundamental challenge to optimize rewards while also discovering as many diverse strategies as possible, which can be crucial in many practical applications. Our study examines two design choices for tackling this challenge, i.e., diversity measure and computation framework. First, we find that with existing diversity measures, visually indistinguishable policies can still yield high diversity scores. To accurately capture the behavioral difference, we propose to incorporate the state-space distance information into the diversity measure. In addition, we examine two common computation frameworks for this problem, i.e., population-based training (PBT) and iterative learning (ITR). We show that although PBT is the precise problem formulation, ITR can achieve comparable diversity scores with higher computation efficiency, leading to improved solution quality in practice. Based on our analysis, we further combine ITR with two tractable realizations of the state-distance-based diversity measures and develop a novel diversity-driven RL algorithm, State-based Intrinsic-reward Policy Optimization (SIPO), with provable convergence properties. We empirically examine SIPO across three domains from robot locomotion to multi-agent games. In all of our testing environments, SIPO consistently produces strategically diverse and human-interpretable policies that cannot be discovered by existing baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- SUB-PLAY: Adversarial Policies against Partially Observed Multi-Agent Reinforcement Learning SystemsOubo Ma, Yuwen Pu, Linkang Du, Yang Dai 等CCS 2024 · 被引用 6 次
- Discovering Creative Behaviors through DUPLEX: Diverse Universal Features for Policy ExplorationBorja G. León, Francesco Riccio, Kaushik Subramanian, Peter R. Wurman 等NeurIPS 2024 · 被引用 5 次
- Moving Out: Physically-grounded Human-AI CollaborationXuhui Kang, Sung-Wook Lee, Haolin Liu, Yuyan Wang 等ICML 2026
它引用的顶会 Paper27
- Emergent Tool Use From Multi-Agent AutocurriculaBowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu 等ICLR 2020 · 被引用 751 次
- On Gradient Descent Ascent for Nonconvex-Concave Minimax ProblemsTianyi Lin, Chi Jin, Michael I. JordanICML 2020 · 被引用 587 次
- Google Research Football: A Novel Reinforcement Learning EnvironmentKarol Kurach, Anton Raichuk, Piotr Stanczyk, Michal Zajac 等AAAI 2020 · 被引用 496 次
- Behavior From the Void: Unsupervised Active Pre-TrainingHao Liu, Pieter AbbeelNeurIPS 2021 · 被引用 258 次
- Count-Based Exploration with the Successor RepresentationMarlos C. Machado, Marc G. Bellemare, Michael BowlingAAAI 2020 · 被引用 206 次
相关 Paper
- Continuously Discovering Novel Strategies via Reward-Switching Policy OptimizationZihan Zhou, Wei Fu, Bingliang Zhang, Yi WuICLR 2022 · 被引用 34 次
- DGPO: Discovering Multiple Strategies with Diversity-Guided Policy OptimizationWentse Chen, Shiyu Huang, Yuan Chiang, Tim Pearce 等AAAI 2024 · 被引用 9 次
- Discovering Policies with DOMiNO: Diversity Optimization Maintaining Near OptimalityTom Zahavy, Yannick Schroecker, Feryal M. P. Behbahani, Kate Baumli 等ICLR 2023 · 被引用 2 次
- Generating Diverse Cooperative Agents by Learning Incompatible PoliciesRujikorn Charakorn, Poramate Manoonpong, Nat DilokthanakulICLR 2023
- Learning Diverse Risk Preferences in Population-Based Self-PlayYuhua Jiang, Qihan Liu, Xiaoteng Ma, Chenghao Li 等AAAI 2024 · 被引用 8 次
