Learning to Play No-Press Diplomacy with Best Response Policy Iteration
Thomas W. Anthony, Tom Eccles, Andrea Tacchetti, János Kramár, Ian Gemp, Thomas C. Hudson, Nicolas Porcel, Marc Lanctot, Julien Pérolat, Richard Everett, Satinder Singh, Thore Graepel, Yoram Bachrach
Abstract
Recent advances in deep reinforcement learning (RL) have led to considerable progress in many 2-player zero-sum games, such as Go, Poker and Starcraft. The purely adversarial nature of such games allows for conceptually simple and principled application of RL methods. However real-world settings are many-agent, and agent interactions are complex mixtures of common-interest and competitive aspects. We consider Diplomacy, a 7-player board game designed to accentuate dilemmas resulting from many-agent interactions. It also features a large combinatorial action space and simultaneous moves, which are challenging for RL algorithms. We propose a simple yet effective approximate best response operator, designed to handle large combinatorial action spaces and simultaneous moves. We also introduce a family of policy iteration methods that approximate fictitious play. With these methods, we successfully apply RL to Diplomacy: we show that our agents convincingly outperform the previous state-of-the-art, and game theoretic equilibrium analysis shows that the new process yields consistent improvements.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 699f8ec5-9e23-4370-8ecc-0f0afc6cffc7Cited by top-tier papers13
- Modeling Strong and Human-Like Gameplay with KL-Regularized SearchAthul Paul Jacob, David J. Wu, Gabriele Farina, Adam Lerer et al.ICML 2022 · 69 citations
- Human-Level Performance in No-Press Diplomacy via Equilibrium SearchJonathan Gray, Adam Lerer, Anton Bakhtin, Noam BrownICLR 2021 · 61 citations
- No-Press Diplomacy from ScratchAnton Bakhtin, David J. Wu, Adam Lerer, Noam BrownNeurIPS 2021 · 51 citations
- Richelieu: Self-Evolving LLM-Based Agents for AI DiplomacyZhenyu Guan, Xiangyu Kong, Fangwei Zhong, Yizhou WangNeurIPS 2024 · 48 citations
- Multi-Agent Training beyond Zero-Sum with Correlated Equilibrium Meta-SolversLuke Marris, Paul Muller, Marc Lanctot, Karl Tuyls et al.ICML 2021 · 42 citations
Builds on7
- Emergent Tool Use From Multi-Agent AutocurriculaBowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu et al.ICLR 2020 · 751 citations
- From Poincaré Recurrence to Convergence in Imperfect Information Games: Finding Equilibrium via RegularizationJulien Pérolat, Rémi Munos, Jean-Baptiste Lespiau, Shayegan Omidshafiei et al.ICML 2021 · 102 citations
- Simplified Action Decoder for Deep Multi-Agent Reinforcement LearningHengyuan Hu, Jakob N. FoersterICLR 2020 · 88 citations
- Improving Policies via Search in Cooperative Partially Observable GamesAdam Lerer, Hengyuan Hu, Jakob N. Foerster, Noam BrownAAAI 2020 · 87 citations
- Arena: A General Evaluation Platform and Building Toolkit for Multi-Agent IntelligenceYuhang Song, Andrzej Wojcicki, Thomas Lukasiewicz, Jianyi Wang et al.AAAI 2020 · 36 citations
Related papers
- Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and PlanningAnton Bakhtin, David J. Wu, Adam Lerer, Jonathan Gray et al.ICLR 2023 · 10 citations
- A Generalized Training Approach for Multiagent LearningPaul Muller, Shayegan Omidshafiei, Mark Rowland, Karl Tuyls et al.ICLR 2020 · 110 citations
- DipLLM: Fine-Tuning LLM for Strategic Decision-making in DiplomacyKaixuan Xu, Jiajun Chai, Sicheng Li, Yuqian Fu et al.ICML 2025
- More Victories, Less Cooperation: Assessing Cicero's Diplomacy PlayWichayaporn Wongkamjan, Feng Gu, Yanze Wang, Ulf Hermjakob et al.ACL 2024
- Explicit Exploration for High-Welfare Equilibria in Game-Theoretic Multiagent Reinforcement LearningAustin A. Nguyen, Anri Gu, Michael P. WellmanICML 2025
