Trajectory Diversity for Zero-Shot Coordination
Andrei Lupu, Brandon Cui, Hengyuan Hu, Jakob N. Foerster
Abstract
We study the problem of zero-shot coordination (ZSC), where agents must independently produce strategies for a collaborative game that are compatible with novel partners not seen during training. Our first contribution is to consider the need for diversity in generating such agents. Because self-play (SP) agents control their own trajectory distribution during training, each policy typically only performs well on this exact distribution. As a result, they achieve low scores in ZSC, since playing with another agent is likely to put them in situations they have not encountered during training. To address this issue, we train a common best response (BR) to a population of agents, which we regulate to be diverse. To this end, we introduce Trajectory Diversity (TrajeDi) – a differentiable objective for generating diverse reinforcement learning policies. We derive TrajeDi as a generalization of the Jensen-Shannon divergence between policies and motivate it experimentally in two simple settings. We then focus on the collaborative card game Hanabi, demonstrating the scalability of our method and improving upon the cross-play scores of both independently trained SP agents and BRs to unregularized populations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c46fe2f8-9e24-4171-bf59-76a3c1743070Cited by top-tier papers54
- Collaborating with Humans without Human DataDJ Strouse, Kevin R. McKee, Matt M. Botvinick, Edward Hughes et al.NeurIPS 2021 · 239 citations
- ProAgent: Building Proactive Cooperative Agents with Large Language ModelsCeyao Zhang, Kaijie Yang, Siyi Hu, Zihao Wang et al.AAAI 2024 · 141 citations
- Maximum Entropy Population-Based Training for Zero-Shot Human-AI CoordinationRui Zhao, Jinming Song, Yufeng Yuan, Haifeng Hu et al.AAAI 2023 · 94 citations
- Language Instructed Reinforcement Learning for Human-AI CoordinationHengyuan Hu, Dorsa SadighICML 2023 · 90 citations
- Revisiting Some Common Practices in Cooperative Multi-Agent Reinforcement LearningWei Fu, Chao Yu, Zelai Xu, Jiaqi Yang et al.ICML 2022 · 49 citations
Builds on5
- Adversarial Policies: Attacking Deep Reinforcement LearningAdam Gleave, Michael Dennis, Cody Wild, Neel Kant et al.ICLR 2020 · 415 citations
- "Other-Play" for Zero-Shot CoordinationHengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob N. FoersterICML 2020 · 271 citations
- Effective Diversity in Population Based Reinforcement LearningJack Parker-Holder, Aldo Pacchiano, Krzysztof Marcin Choromanski, Stephen J. RobertsNeurIPS 2020 · 195 citations
- One Solution is Not All You Need: Few-Shot Extrapolation via Structured MaxEnt RLSaurabh Kumar, Aviral Kumar, Sergey Levine, Chelsea FinnNeurIPS 2020 · 109 citations
- Discovering Diverse Multi-Agent Strategic Behavior via Reward RandomizationZhenggang Tang, Chao Yu, Boyuan Chen, Huazhe Xu et al.ICLR 2021 · 63 citations
Related papers
- Efficient Reinforcement Learning for Zero-Shot Coordination in Evolving GamesBingyu Hui, Lebin Yu, Quanming Yao, Yunpeng Qu et al.AAAI 2026 · 1 citation
- Off-Team LearningBrandon Cui, Hengyuan Hu, Andrei Lupu, Samuel Sokota et al.NeurIPS 2022 · 4 citations
- Adaptive Coordination in Social Embodied RearrangementAndrew Szot, Unnat Jain, Dhruv Batra, Zsolt Kira et al.ICML 2023 · 20 citations
- K-level Reasoning for Zero-Shot Coordination in HanabiBrandon Cui, Hengyuan Hu, Luis Pineda, Jakob N. FoersterNeurIPS 2021 · 46 citations
- Adversarial Diversity in HanabiBrandon Cui, Andrei Lupu, Samuel Sokota, Hengyuan Hu et al.ICLR 2023
