Generalized Beliefs for Cooperative AI
Darius Muglich, Luisa M. Zintgraf, Christian A. Schröder de Witt, Shimon Whiteson, Jakob N. Foerster
摘要
Self-play is a common paradigm for constructing solutions in Markov games that can yield optimal policies in collaborative settings. However, these policies often adopt highly-specialized conventions that make playing with a novel partner difficult. To address this, recent approaches rely on encoding symmetry and convention-awareness into policy training, but these require strong environmental assumptions and can complicate policy training. We therefore propose moving the learning of conventions to the belief space. Specifically, we propose a belief learning model that can maintain beliefs over rollouts of policies not seen at training time, and can thus decode and adapt to novel conventions at test time. We show how to leverage this model for both search and training of a best response over various pools of policies to greatly improve ad-hoc teamplay. We also show how our setup promotes explainability and interpretability of nuanced agent conventions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Secret Collusion among AI Agents: Multi-Agent Deception via SteganographySumeet Ramesh Motwani, Mikhail Baranchuk, Martin Strohmeier, Vijay Bolina 等NeurIPS 2024 · 被引用 140 次
- Equivariant Networks for Zero-Shot CoordinationDarius Muglich, Christian Schröder de Witt, Elise van der Pol, Shimon Whiteson 等NeurIPS 2022 · 被引用 24 次
- Diversity Is Not All You Need: Training A Robust Cooperative Agent Needs Specialist PartnersRujikorn Charakorn, Poramate Manoonpong, Nat DilokthanakulNeurIPS 2024 · 被引用 3 次
- GlobeDiff: State Diffusion Process for Partial Observability in Multi-Agent SystemYiqin Yang, Xu Yang, Yuhua Jiang, Ni Mu 等ICLR 2026
- Expected Return SymmetriesDarius Muglich, Johannes Forkel, Elise van der Pol, Jakob Nicolaus FoersterICLR 2025
它引用的顶会 Paper11
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Stabilizing Transformers for Reinforcement LearningEmilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu 等ICML 2020 · 被引用 464 次
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-LearningLuisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze 等ICLR 2020 · 被引用 315 次
- "Other-Play" for Zero-Shot CoordinationHengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob N. FoersterICML 2020 · 被引用 271 次
- Combining Deep Reinforcement Learning and Search for Imperfect-Information GamesNoam Brown, Anton Bakhtin, Adam Lerer, Qucheng GongNeurIPS 2020 · 被引用 205 次
相关 Paper
- Off-Belief LearningHengyuan Hu, Adam Lerer, Brandon Cui, Luis Pineda 等ICML 2021 · 被引用 86 次
- Diverse Conventions for Human-AI CollaborationBidipta Sarkar, Andy Shih, Dorsa SadighNeurIPS 2023 · 被引用 23 次
- K-level Reasoning for Zero-Shot Coordination in HanabiBrandon Cui, Hengyuan Hu, Luis Pineda, Jakob N. FoersterNeurIPS 2021 · 被引用 46 次
- On the Critical Role of Conventions in Adaptive Human-AI CollaborationAndy Shih, Arjun Sawhney, Jovana Kondic, Stefano Ermon 等ICLR 2021 · 被引用 46 次
- Back to the Future: Toward a Hybrid Architecture for Ad Hoc TeamworkHasra Dodampegama, Mohan SridharanAAAI 2023 · 被引用 8 次
