Equivariant Networks for Zero-Shot Coordination
Darius Muglich, Christian Schröder de Witt, Elise van der Pol, Shimon Whiteson, Jakob N. Foerster
摘要
Successful coordination in Dec-POMDPs requires agents to adopt robust strategies and interpretable styles of play for their partner. A common failure mode is symmetry breaking, when agents arbitrarily converge on one out of many equivalent but mutually incompatible policies. Commonly these examples include partial observability, e.g. waving your right hand vs. left hand to convey a covert message. In this paper, we present a novel equivariant network architecture for use in Dec-POMDPs that effectively leverages environmental symmetry for improving zero-shot coordination, doing so more effectively than prior methods. Our method also acts as a ``coordination-improvement operator'' for generic, pre-trained policies, and thus may be applied at test-time in conjunction with any self-play algorithm. We provide theoretical guarantees of our work and test on the AI benchmark task of Hanabi, where we demonstrate our methods outperforming other symmetry-aware baselines in zero-shot coordination, as well as able to improve the coordination ability of a variety of pre-trained policies. In particular, we show our method can be used to improve on the state of the art for zero-shot coordination on the Hanabi benchmark.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Secret Collusion among AI Agents: Multi-Agent Deception via SteganographySumeet Ramesh Motwani, Mikhail Baranchuk, Martin Strohmeier, Vijay Bolina 等NeurIPS 2024 · 被引用 140 次
- EDGI: Equivariant Diffusion for Planning with Embodied AgentsJohann Brehmer, Joey Bose, Pim de Haan, Taco S. CohenNeurIPS 2023 · 被引用 52 次
- Efficient Equivariant Transfer Learning from Pretrained ModelsSourya Basu, Pulkit Katdare, Prasanna Sattigeri, Vijil Chenthamarakshan 等NeurIPS 2023 · 被引用 14 次
- Who Needs to Know? Minimal Knowledge for Optimal CoordinationNiklas Lauffer, Ameesh Shah, Micah Carroll, Michael D. Dennis 等ICML 2023 · 被引用 8 次
- Robust and Diverse Multi-Agent Learning via Rational Policy GradientNiklas Lauffer, Ameesh Shah, Micah Carroll, Sanjit A. Seshia 等NeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper16
- Graph Convolutional Reinforcement LearningJiechuan Jiang, Chen Dun, Tiejun Huang, Zongqing LuICLR 2020 · 被引用 415 次
- "Other-Play" for Zero-Shot CoordinationHengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob N. FoersterICML 2020 · 被引用 271 次
- A Practical Method for Constructing Equivariant Multilayer Perceptrons for Arbitrary Matrix GroupsMarc Finzi, Max Welling, Andrew Gordon WilsonICML 2021 · 被引用 226 次
- Deep Coordination GraphsWendelin Boehmer, Vitaly Kurin, Shimon WhitesonICML 2020 · 被引用 209 次
- MDP Homomorphic Networks: Group Symmetries in Reinforcement LearningElise van der Pol, Daniel E. Worrall, Herke van Hoof, Frans A. Oliehoek 等NeurIPS 2020 · 被引用 203 次
相关 Paper
- K-level Reasoning for Zero-Shot Coordination in HanabiBrandon Cui, Hengyuan Hu, Luis Pineda, Jakob N. FoersterNeurIPS 2021 · 被引用 46 次
- Expected Return SymmetriesDarius Muglich, Johannes Forkel, Elise van der Pol, Jakob Nicolaus FoersterICLR 2025
- Off-Team LearningBrandon Cui, Hengyuan Hu, Andrei Lupu, Samuel Sokota 等NeurIPS 2022 · 被引用 4 次
- Off-Belief LearningHengyuan Hu, Adam Lerer, Brandon Cui, Luis Pineda 等ICML 2021 · 被引用 86 次
- Adversarial Diversity in HanabiBrandon Cui, Andrei Lupu, Samuel Sokota, Hengyuan Hu 等ICLR 2023
