Learning Diverse Risk Preferences in Population-Based Self-Play
Yuhua Jiang, Qihan Liu, Xiaoteng Ma, Chenghao Li, Yiqin Yang, Jun Yang, Bin Liang, Qianchuan Zhao
摘要
Among the remarkable successes of Reinforcement Learning (RL), self-play algorithms have played a crucial role in solving competitive games. However, current self-play RL methods commonly optimize the agent to maximize the expected win-rates against its current or historical copies, resulting in a limited strategy style and a tendency to get stuck in local optima. To address this limitation, it is important to improve the diversity of policies, allowing the agent to break stalemates and enhance its robustness when facing with different opponents. In this paper, we present a novel perspective to promote diversity by considering that agents could have diverse risk preferences in the face of uncertainty. To achieve this, we introduce a novel reinforcement learning algorithm called Risk-sensitive Proximal Policy Optimization (RPPO), which smoothly interpolates between worst-case and best-case policy learning, enabling policy learning with desired risk preferences. Furthermore, by seamlessly integrating RPPO with population-based self-play, agents in the population optimize dynamic risk-sensitive objectives using experiences gained from playing against diverse opponents. Our empirical results demonstrate that our method achieves comparable or superior performance in competitive games and, importantly, leads to the emergence of diverse behavioral modes. Code is available at https://github.com/Jackory/RPBT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Risk-Sensitive Reinforcement Learning for Alleviating Exploration Dilemmas in Large Language ModelsYuhua Jiang, Jiawei Huang, Yufeng Yuan, Xin Mao 等ICLR 2026 · 被引用 8 次
- Craftium: Bridging Flexibility and Efficiency for Rich 3D Single- and Multi-Agent EnvironmentsMikel Malagón, Josu Ceberio, José Antonio LozanoICML 2025
- Episodic Novelty Through Temporal DistanceYuhua Jiang, Qihan Liu, Yiqin Yang, Xiaoteng Ma 等ICLR 2025
它引用的顶会 Paper15
- Emergent Tool Use From Multi-Agent AutocurriculaBowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu 等ICLR 2020 · 被引用 751 次
- Collaborating with Humans without Human DataDJ Strouse, Kevin R. McKee, Matt M. Botvinick, Edward Hughes 等NeurIPS 2021 · 被引用 239 次
- Effective Diversity in Population Based Reinforcement LearningJack Parker-Holder, Aldo Pacchiano, Krzysztof Marcin Choromanski, Stephen J. RobertsNeurIPS 2020 · 被引用 195 次
- WCSAC: Worst-Case Soft Actor Critic for Safety-Constrained Reinforcement LearningQisong Yang, Thiago D. Simão, Simon H. Tindemans, Matthijs T. J. SpaanAAAI 2021 · 被引用 168 次
- Trajectory Diversity for Zero-Shot CoordinationAndrei Lupu, Brandon Cui, Hengyuan Hu, Jakob N. FoersterICML 2021 · 被引用 157 次
相关 Paper
- Proximal Policy Gradient Arborescence for Quality Diversity Reinforcement LearningSumeet Batra, Bryon Tjanaka, Matthew Christopher Fontaine, Aleksei Petrenko 等ICLR 2024 · 被引用 26 次
- Coordinated Proximal Policy OptimizationZifan Wu, Chao Yu, Deheng Ye, Junge Zhang 等NeurIPS 2021 · 被引用 73 次
- Adversarial Policy Training against Deep Reinforcement LearningXian Wu, Wenbo Guo, Hua Wei, Xinyu XingUSENIX Security 2021 · 被引用 19 次
- Polychromic Objectives for Reinforcement LearningJubayer Ibn Hamid, Ifdita Hasan Orney, Ellen Xu, Chelsea Finn 等ICLR 2026 · 被引用 9 次
- Rethinking the Trust Region in LLM Reinforcement LearningPenghui Qi, Xiangxin Zhou, Zichen Liu, Tianyu Pang 等ICML 2026 · 被引用 22 次
