Generalized Beliefs for Cooperative AI
Darius Muglich, Luisa M. Zintgraf, Christian A. Schröder de Witt, Shimon Whiteson, Jakob N. Foerster
Abstract
Self-play is a common paradigm for constructing solutions in Markov games that can yield optimal policies in collaborative settings. However, these policies often adopt highly-specialized conventions that make playing with a novel partner difficult. To address this, recent approaches rely on encoding symmetry and convention-awareness into policy training, but these require strong environmental assumptions and can complicate policy training. We therefore propose moving the learning of conventions to the belief space. Specifically, we propose a belief learning model that can maintain beliefs over rollouts of policies not seen at training time, and can thus decode and adapt to novel conventions at test time. We show how to leverage this model for both search and training of a best response over various pools of policies to greatly improve ad-hoc teamplay. We also show how our setup promotes explainability and interpretability of nuanced agent conventions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Secret Collusion among AI Agents: Multi-Agent Deception via SteganographySumeet Ramesh Motwani, Mikhail Baranchuk, Martin Strohmeier, Vijay Bolina et al.NeurIPS 2024 · 140 citations
- Equivariant Networks for Zero-Shot CoordinationDarius Muglich, Christian Schröder de Witt, Elise van der Pol, Shimon Whiteson et al.NeurIPS 2022 · 24 citations
- Diversity Is Not All You Need: Training A Robust Cooperative Agent Needs Specialist PartnersRujikorn Charakorn, Poramate Manoonpong, Nat DilokthanakulNeurIPS 2024 · 3 citations
- GlobeDiff: State Diffusion Process for Partial Observability in Multi-Agent SystemYiqin Yang, Xu Yang, Yuhua Jiang, Ni Mu et al.ICLR 2026
- Expected Return SymmetriesDarius Muglich, Johannes Forkel, Elise van der Pol, Jakob Nicolaus FoersterICLR 2025
Builds on11
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Stabilizing Transformers for Reinforcement LearningEmilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu et al.ICML 2020 · 464 citations
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-LearningLuisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze et al.ICLR 2020 · 315 citations
- "Other-Play" for Zero-Shot CoordinationHengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob N. FoersterICML 2020 · 271 citations
- Combining Deep Reinforcement Learning and Search for Imperfect-Information GamesNoam Brown, Anton Bakhtin, Adam Lerer, Qucheng GongNeurIPS 2020 · 205 citations
Related papers
- Off-Belief LearningHengyuan Hu, Adam Lerer, Brandon Cui, Luis Pineda et al.ICML 2021 · 86 citations
- Diverse Conventions for Human-AI CollaborationBidipta Sarkar, Andy Shih, Dorsa SadighNeurIPS 2023 · 23 citations
- K-level Reasoning for Zero-Shot Coordination in HanabiBrandon Cui, Hengyuan Hu, Luis Pineda, Jakob N. FoersterNeurIPS 2021 · 46 citations
- On the Critical Role of Conventions in Adaptive Human-AI CollaborationAndy Shih, Arjun Sawhney, Jovana Kondic, Stefano Ermon et al.ICLR 2021 · 46 citations
- Back to the Future: Toward a Hybrid Architecture for Ad Hoc TeamworkHasra Dodampegama, Mohan SridharanAAAI 2023 · 8 citations
