Trajectory Diversity for Zero-Shot Coordination
Andrei Lupu, Brandon Cui, Hengyuan Hu, Jakob N. Foerster
摘要
We study the problem of zero-shot coordination (ZSC), where agents must independently produce strategies for a collaborative game that are compatible with novel partners not seen during training. Our first contribution is to consider the need for diversity in generating such agents. Because self-play (SP) agents control their own trajectory distribution during training, each policy typically only performs well on this exact distribution. As a result, they achieve low scores in ZSC, since playing with another agent is likely to put them in situations they have not encountered during training. To address this issue, we train a common best response (BR) to a population of agents, which we regulate to be diverse. To this end, we introduce Trajectory Diversity (TrajeDi) – a differentiable objective for generating diverse reinforcement learning policies. We derive TrajeDi as a generalization of the Jensen-Shannon divergence between policies and motivate it experimentally in two simple settings. We then focus on the collaborative card game Hanabi, demonstrating the scalability of our method and improving upon the cross-play scores of both independently trained SP agents and BRs to unregularized populations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper54
- Collaborating with Humans without Human DataDJ Strouse, Kevin R. McKee, Matt M. Botvinick, Edward Hughes 等NeurIPS 2021 · 被引用 239 次
- ProAgent: Building Proactive Cooperative Agents with Large Language ModelsCeyao Zhang, Kaijie Yang, Siyi Hu, Zihao Wang 等AAAI 2024 · 被引用 141 次
- Maximum Entropy Population-Based Training for Zero-Shot Human-AI CoordinationRui Zhao, Jinming Song, Yufeng Yuan, Haifeng Hu 等AAAI 2023 · 被引用 94 次
- Language Instructed Reinforcement Learning for Human-AI CoordinationHengyuan Hu, Dorsa SadighICML 2023 · 被引用 90 次
- Revisiting Some Common Practices in Cooperative Multi-Agent Reinforcement LearningWei Fu, Chao Yu, Zelai Xu, Jiaqi Yang 等ICML 2022 · 被引用 49 次
它引用的顶会 Paper5
- Adversarial Policies: Attacking Deep Reinforcement LearningAdam Gleave, Michael Dennis, Cody Wild, Neel Kant 等ICLR 2020 · 被引用 415 次
- "Other-Play" for Zero-Shot CoordinationHengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob N. FoersterICML 2020 · 被引用 271 次
- Effective Diversity in Population Based Reinforcement LearningJack Parker-Holder, Aldo Pacchiano, Krzysztof Marcin Choromanski, Stephen J. RobertsNeurIPS 2020 · 被引用 195 次
- One Solution is Not All You Need: Few-Shot Extrapolation via Structured MaxEnt RLSaurabh Kumar, Aviral Kumar, Sergey Levine, Chelsea FinnNeurIPS 2020 · 被引用 109 次
- Discovering Diverse Multi-Agent Strategic Behavior via Reward RandomizationZhenggang Tang, Chao Yu, Boyuan Chen, Huazhe Xu 等ICLR 2021 · 被引用 63 次
相关 Paper
- Efficient Reinforcement Learning for Zero-Shot Coordination in Evolving GamesBingyu Hui, Lebin Yu, Quanming Yao, Yunpeng Qu 等AAAI 2026 · 被引用 1 次
- Off-Team LearningBrandon Cui, Hengyuan Hu, Andrei Lupu, Samuel Sokota 等NeurIPS 2022 · 被引用 4 次
- Adaptive Coordination in Social Embodied RearrangementAndrew Szot, Unnat Jain, Dhruv Batra, Zsolt Kira 等ICML 2023 · 被引用 20 次
- K-level Reasoning for Zero-Shot Coordination in HanabiBrandon Cui, Hengyuan Hu, Luis Pineda, Jakob N. FoersterNeurIPS 2021 · 被引用 46 次
- Adversarial Diversity in HanabiBrandon Cui, Andrei Lupu, Samuel Sokota, Hengyuan Hu 等ICLR 2023
