K-level Reasoning for Zero-Shot Coordination in Hanabi
Brandon Cui, Hengyuan Hu, Luis Pineda, Jakob N. Foerster
摘要
The standard problem setting in cooperative multi-agent settings is self-play (SP), where the goal is to train a team of agents that works well together. However, optimal SP policies commonly contain arbitrary conventions ("handshakes") and are not compatible with other, independently trained agents or humans. This latter desiderata was recently formalized by Hu et al. 2020 as the zero-shot coordination (ZSC) setting and partially addressed with their Other-Play (OP) algorithm, which showed improved ZSC and human-AI performance in the card game Hanabi. OP assumes access to the symmetries of the environment and prevents agents from breaking these in a mutually incompatible way during training. However, as the authors point out, discovering symmetries for a given environment is a computationally hard problem. Instead, we show that through a simple adaption of k-level reasoning (KLR) Costa Gomes et al. 2006, synchronously training all levels, we can obtain competitive ZSC and ad-hoc teamplay performance in Hanabi, including when paired with a human-like proxy bot. We also introduce a new method, synchronous-k-level reasoning with a best response (SyKLRBR), which further improves performance on our synchronous KLR by co-training a best response.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Language Instructed Reinforcement Learning for Human-AI CoordinationHengyuan Hu, Dorsa SadighICML 2023 · 被引用 90 次
- Equivariant Networks for Zero-Shot CoordinationDarius Muglich, Christian Schröder de Witt, Elise van der Pol, Shimon Whiteson 等NeurIPS 2022 · 被引用 24 次
- Sample-Efficient Quality-Diversity by Cooperative CoevolutionKe Xue, Ren-Jian Wang, Pengyi Li, Dong Li 等ICLR 2024 · 被引用 17 次
- Learning General World Models in a Handful of Reward-Free DeploymentsYingchen Xu, Jack Parker-Holder, Aldo Pacchiano, Philip J. Ball 等NeurIPS 2022 · 被引用 16 次
- On the Impossibility of Learning to Cooperate with Adaptive Partner Strategies in Repeated GamesRobert Tyler Loftin, Frans A. OliehoekICML 2022 · 被引用 4 次
它引用的顶会 Paper6
- "Other-Play" for Zero-Shot CoordinationHengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob N. FoersterICML 2020 · 被引用 271 次
- Simplified Action Decoder for Deep Multi-Agent Reinforcement LearningHengyuan Hu, Jakob N. FoersterICLR 2020 · 被引用 88 次
- Off-Belief LearningHengyuan Hu, Adam Lerer, Brandon Cui, Luis Pineda 等ICML 2021 · 被引用 86 次
- Continuous Coordination As a Realistic Scenario for Lifelong LearningHadi Nekoei, Akilesh Badrinaaraayanan, Aaron C. Courville, Sarath ChandarICML 2021 · 被引用 51 次
- SEED RL: Scalable and Efficient Deep-RL with Accelerated Central InferenceLasse Espeholt, Raphaël Marinier, Piotr Stanczyk, Ke Wang 等ICLR 2020 · 被引用 32 次
相关 Paper
- A New Formalism, Method and Open Issues for Zero-Shot CoordinationJohannes Treutlein, Michael Dennis, Caspar Oesterheld, Jakob N. FoersterICML 2021 · 被引用 45 次
- Off-Team LearningBrandon Cui, Hengyuan Hu, Andrei Lupu, Samuel Sokota 等NeurIPS 2022 · 被引用 4 次
- Improving Policies via Search in Cooperative Partially Observable GamesAdam Lerer, Hengyuan Hu, Jakob N. Foerster, Noam BrownAAAI 2020 · 被引用 87 次
- Trajectory Diversity for Zero-Shot CoordinationAndrei Lupu, Brandon Cui, Hengyuan Hu, Jakob N. FoersterICML 2021 · 被引用 157 次
- The Hidden Rules of Hanabi: How Humans Outperform AI AgentsMatthew Sidji, Wally Smith, Melissa J. RogersonCHI 2023 · 被引用 9 次
