Equivariant Networks for Zero-Shot Coordination
Darius Muglich, Christian Schröder de Witt, Elise van der Pol, Shimon Whiteson, Jakob N. Foerster
Abstract
Successful coordination in Dec-POMDPs requires agents to adopt robust strategies and interpretable styles of play for their partner. A common failure mode is symmetry breaking, when agents arbitrarily converge on one out of many equivalent but mutually incompatible policies. Commonly these examples include partial observability, e.g. waving your right hand vs. left hand to convey a covert message. In this paper, we present a novel equivariant network architecture for use in Dec-POMDPs that effectively leverages environmental symmetry for improving zero-shot coordination, doing so more effectively than prior methods. Our method also acts as a ``coordination-improvement operator'' for generic, pre-trained policies, and thus may be applied at test-time in conjunction with any self-play algorithm. We provide theoretical guarantees of our work and test on the AI benchmark task of Hanabi, where we demonstrate our methods outperforming other symmetry-aware baselines in zero-shot coordination, as well as able to improve the coordination ability of a variety of pre-trained policies. In particular, we show our method can be used to improve on the state of the art for zero-shot coordination on the Hanabi benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ae9fabe-95a7-4ba2-b9e8-3e3c3faf0fd3Cited by top-tier papers6
- Secret Collusion among AI Agents: Multi-Agent Deception via SteganographySumeet Ramesh Motwani, Mikhail Baranchuk, Martin Strohmeier, Vijay Bolina et al.NeurIPS 2024 · 140 citations
- EDGI: Equivariant Diffusion for Planning with Embodied AgentsJohann Brehmer, Joey Bose, Pim de Haan, Taco S. CohenNeurIPS 2023 · 52 citations
- Efficient Equivariant Transfer Learning from Pretrained ModelsSourya Basu, Pulkit Katdare, Prasanna Sattigeri, Vijil Chenthamarakshan et al.NeurIPS 2023 · 14 citations
- Who Needs to Know? Minimal Knowledge for Optimal CoordinationNiklas Lauffer, Ameesh Shah, Micah Carroll, Michael D. Dennis et al.ICML 2023 · 8 citations
- Robust and Diverse Multi-Agent Learning via Rational Policy GradientNiklas Lauffer, Ameesh Shah, Micah Carroll, Sanjit A. Seshia et al.NeurIPS 2025 · 4 citations
Builds on16
- Graph Convolutional Reinforcement LearningJiechuan Jiang, Chen Dun, Tiejun Huang, Zongqing LuICLR 2020 · 415 citations
- "Other-Play" for Zero-Shot CoordinationHengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob N. FoersterICML 2020 · 271 citations
- A Practical Method for Constructing Equivariant Multilayer Perceptrons for Arbitrary Matrix GroupsMarc Finzi, Max Welling, Andrew Gordon WilsonICML 2021 · 226 citations
- Deep Coordination GraphsWendelin Boehmer, Vitaly Kurin, Shimon WhitesonICML 2020 · 209 citations
- MDP Homomorphic Networks: Group Symmetries in Reinforcement LearningElise van der Pol, Daniel E. Worrall, Herke van Hoof, Frans A. Oliehoek et al.NeurIPS 2020 · 203 citations
Related papers
- K-level Reasoning for Zero-Shot Coordination in HanabiBrandon Cui, Hengyuan Hu, Luis Pineda, Jakob N. FoersterNeurIPS 2021 · 46 citations
- Expected Return SymmetriesDarius Muglich, Johannes Forkel, Elise van der Pol, Jakob Nicolaus FoersterICLR 2025
- Off-Team LearningBrandon Cui, Hengyuan Hu, Andrei Lupu, Samuel Sokota et al.NeurIPS 2022 · 4 citations
- Off-Belief LearningHengyuan Hu, Adam Lerer, Brandon Cui, Luis Pineda et al.ICML 2021 · 86 citations
- Adversarial Diversity in HanabiBrandon Cui, Andrei Lupu, Samuel Sokota, Hengyuan Hu et al.ICLR 2023
