Off-Belief Learning
Hengyuan Hu, Adam Lerer, Brandon Cui, Luis Pineda, Noam Brown, Jakob N. Foerster
Abstract
The standard problem setting in Dec-POMDPs is self-play, where the goal is to find a set of policies that play optimally together. Policies learned through self-play may adopt arbitrary conventions and implicitly rely on multi-step reasoning based on fragile assumptions about other agents' actions and thus fail when paired with humans or independently trained agents at test time. To address this, we present off-belief learning (OBL). At each timestep OBL agents follow a policy that is optimized assuming past actions were taken by a given, fixed policy (), but assuming that future actions will be taken by . When is uniform random, OBL converges to an optimal policy that does not rely on inferences based on other agents' behavior (an optimal grounded policy). OBL can be iterated in a hierarchy, where the optimal policy from one level becomes the input to the next, thereby introducing multi-level cognitive reasoning in a controlled manner. Unlike existing approaches, which may converge to any equilibrium policy, OBL converges to a unique policy, making it suitable for zero-shot coordination (ZSC). OBL can be scaled to high-dimensional settings with a fictitious transition mechanism and shows strong performance in both a toy-setting and the benchmark human-AI&ZSC problem Hanabi.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aaf46fe7-9477-419f-a352-5908b751555aCited by top-tier papers27
- Secret Collusion among AI Agents: Multi-Agent Deception via SteganographySumeet Ramesh Motwani, Mikhail Baranchuk, Martin Strohmeier, Vijay Bolina et al.NeurIPS 2024 · 140 citations
- Language Instructed Reinforcement Learning for Human-AI CoordinationHengyuan Hu, Dorsa SadighICML 2023 · 90 citations
- Modeling Strong and Human-Like Gameplay with KL-Regularized SearchAthul Paul Jacob, David J. Wu, Gabriele Farina, Adam Lerer et al.ICML 2022 · 69 citations
- K-level Reasoning for Zero-Shot Coordination in HanabiBrandon Cui, Hengyuan Hu, Luis Pineda, Jakob N. FoersterNeurIPS 2021 · 46 citations
- The Boltzmann Policy Distribution: Accounting for Systematic Suboptimality in Human ModelsCassidy Laidlaw, Anca D. DraganICLR 2022 · 46 citations
Builds on3
- "Other-Play" for Zero-Shot CoordinationHengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob N. FoersterICML 2020 · 271 citations
- Simplified Action Decoder for Deep Multi-Agent Reinforcement LearningHengyuan Hu, Jakob N. FoersterICLR 2020 · 88 citations
- SEED RL: Scalable and Efficient Deep-RL with Accelerated Central InferenceLasse Espeholt, Raphaël Marinier, Piotr Stanczyk, Ke Wang et al.ICLR 2020 · 32 citations
Related papers
- Off-Team LearningBrandon Cui, Hengyuan Hu, Andrei Lupu, Samuel Sokota et al.NeurIPS 2022 · 4 citations
- Equivariant Networks for Zero-Shot CoordinationDarius Muglich, Christian Schröder de Witt, Elise van der Pol, Shimon Whiteson et al.NeurIPS 2022 · 24 citations
- Adversarial Diversity in HanabiBrandon Cui, Andrei Lupu, Samuel Sokota, Hengyuan Hu et al.ICLR 2023
- A New Formalism, Method and Open Issues for Zero-Shot CoordinationJohannes Treutlein, Michael Dennis, Caspar Oesterheld, Jakob N. FoersterICML 2021 · 45 citations
- Generalized Beliefs for Cooperative AIDarius Muglich, Luisa M. Zintgraf, Christian A. Schröder de Witt, Shimon Whiteson et al.ICML 2022 · 11 citations
