Off-Team Learning
Brandon Cui, Hengyuan Hu, Andrei Lupu, Samuel Sokota, Jakob N. Foerster
Abstract
Zero-shot coordination (ZSC) evaluates an algorithm by the performance of a team of agents that were trained independently under that algorithm. Off-belief learning (OBL) is a recent method that achieves state-of-the-art results in ZSC in the game Hanabi. However, the implementation of OBL relies on a belief model that experiences covariate shift. Moreover, during ad-hoc coordination, OBL or any other neural policy may experience test-time covariate shift. We present two methods addressing these issues. The first method, off-team belief learning (OT-BL), attempts to improve the accuracy of the belief model of a target policy π T on a broader range of inputs by weighting trajectories approximately according to the distribution induced by a different policy π b . The second, off-team off-belief learning (OT-OBL), attempts to compute an OBL equilibrium, where fixed point error is weighted according to the distribution induced by cross-play between the training policy π and a different fixed policy π b instead of self-play of π . We investigate these methods in variants of Hanabi.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- "Other-Play" for Zero-Shot CoordinationHengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob N. FoersterICML 2020 · 271 citations
- Collaborating with Humans without Human DataDJ Strouse, Kevin R. McKee, Matt M. Botvinick, Edward Hughes et al.NeurIPS 2021 · 239 citations
- Trajectory Diversity for Zero-Shot CoordinationAndrei Lupu, Brandon Cui, Hengyuan Hu, Jakob N. FoersterICML 2021 · 157 citations
- Evaluation of Human-AI Teams for Learned and Rule-Based Agents in HanabiHo Chit Siu, Jaime Daniel Peña, Edenna Chen, Yutai Zhou et al.NeurIPS 2021 · 78 citations
- A New Formalism, Method and Open Issues for Zero-Shot CoordinationJohannes Treutlein, Michael Dennis, Caspar Oesterheld, Jakob N. FoersterICML 2021 · 45 citations
Related papers
- Off-Belief LearningHengyuan Hu, Adam Lerer, Brandon Cui, Luis Pineda et al.ICML 2021 · 86 citations
- Equivariant Networks for Zero-Shot CoordinationDarius Muglich, Christian Schröder de Witt, Elise van der Pol, Shimon Whiteson et al.NeurIPS 2022 · 24 citations
- K-level Reasoning for Zero-Shot Coordination in HanabiBrandon Cui, Hengyuan Hu, Luis Pineda, Jakob N. FoersterNeurIPS 2021 · 46 citations
- Continuous Coordination As a Realistic Scenario for Lifelong LearningHadi Nekoei, Akilesh Badrinaaraayanan, Aaron C. Courville, Sarath ChandarICML 2021 · 51 citations
- Adversarial Diversity in HanabiBrandon Cui, Andrei Lupu, Samuel Sokota, Hengyuan Hu et al.ICLR 2023
