Learning to Play No-Press Diplomacy with Best Response Policy Iteration
Thomas W. Anthony, Tom Eccles, Andrea Tacchetti, János Kramár, Ian Gemp, Thomas C. Hudson, Nicolas Porcel, Marc Lanctot, Julien Pérolat, Richard Everett, Satinder Singh, Thore Graepel, Yoram Bachrach
摘要
Recent advances in deep reinforcement learning (RL) have led to considerable progress in many 2-player zero-sum games, such as Go, Poker and Starcraft. The purely adversarial nature of such games allows for conceptually simple and principled application of RL methods. However real-world settings are many-agent, and agent interactions are complex mixtures of common-interest and competitive aspects. We consider Diplomacy, a 7-player board game designed to accentuate dilemmas resulting from many-agent interactions. It also features a large combinatorial action space and simultaneous moves, which are challenging for RL algorithms. We propose a simple yet effective approximate best response operator, designed to handle large combinatorial action spaces and simultaneous moves. We also introduce a family of policy iteration methods that approximate fictitious play. With these methods, we successfully apply RL to Diplomacy: we show that our agents convincingly outperform the previous state-of-the-art, and game theoretic equilibrium analysis shows that the new process yields consistent improvements.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Modeling Strong and Human-Like Gameplay with KL-Regularized SearchAthul Paul Jacob, David J. Wu, Gabriele Farina, Adam Lerer 等ICML 2022 · 被引用 69 次
- Human-Level Performance in No-Press Diplomacy via Equilibrium SearchJonathan Gray, Adam Lerer, Anton Bakhtin, Noam BrownICLR 2021 · 被引用 61 次
- No-Press Diplomacy from ScratchAnton Bakhtin, David J. Wu, Adam Lerer, Noam BrownNeurIPS 2021 · 被引用 51 次
- Richelieu: Self-Evolving LLM-Based Agents for AI DiplomacyZhenyu Guan, Xiangyu Kong, Fangwei Zhong, Yizhou WangNeurIPS 2024 · 被引用 48 次
- Multi-Agent Training beyond Zero-Sum with Correlated Equilibrium Meta-SolversLuke Marris, Paul Muller, Marc Lanctot, Karl Tuyls 等ICML 2021 · 被引用 42 次
它引用的顶会 Paper7
- Emergent Tool Use From Multi-Agent AutocurriculaBowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu 等ICLR 2020 · 被引用 751 次
- From Poincaré Recurrence to Convergence in Imperfect Information Games: Finding Equilibrium via RegularizationJulien Pérolat, Rémi Munos, Jean-Baptiste Lespiau, Shayegan Omidshafiei 等ICML 2021 · 被引用 102 次
- Simplified Action Decoder for Deep Multi-Agent Reinforcement LearningHengyuan Hu, Jakob N. FoersterICLR 2020 · 被引用 88 次
- Improving Policies via Search in Cooperative Partially Observable GamesAdam Lerer, Hengyuan Hu, Jakob N. Foerster, Noam BrownAAAI 2020 · 被引用 87 次
- Arena: A General Evaluation Platform and Building Toolkit for Multi-Agent IntelligenceYuhang Song, Andrzej Wojcicki, Thomas Lukasiewicz, Jianyi Wang 等AAAI 2020 · 被引用 36 次
相关 Paper
- Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and PlanningAnton Bakhtin, David J. Wu, Adam Lerer, Jonathan Gray 等ICLR 2023 · 被引用 10 次
- A Generalized Training Approach for Multiagent LearningPaul Muller, Shayegan Omidshafiei, Mark Rowland, Karl Tuyls 等ICLR 2020 · 被引用 110 次
- DipLLM: Fine-Tuning LLM for Strategic Decision-making in DiplomacyKaixuan Xu, Jiajun Chai, Sicheng Li, Yuqian Fu 等ICML 2025
- More Victories, Less Cooperation: Assessing Cicero's Diplomacy PlayWichayaporn Wongkamjan, Feng Gu, Yanze Wang, Ulf Hermjakob 等ACL 2024
- Explicit Exploration for High-Welfare Equilibria in Game-Theoretic Multiagent Reinforcement LearningAustin A. Nguyen, Anri Gu, Michael P. WellmanICML 2025
