Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning
Anton Bakhtin, David J. Wu, Adam Lerer, Jonathan Gray, Athul Paul Jacob, Gabriele Farina, Alexander H. Miller, Noam Brown
摘要
No-press Diplomacy is a complex strategy game involving both cooperation and competition that has served as a benchmark for multi-agent AI research. While self-play reinforcement learning has resulted in numerous successes in purely adversarial games like chess, Go, and poker, self-play alone is insufficient for achieving optimal performance in domains involving cooperation with humans. We address this shortcoming by first introducing a planning algorithm we call DiL-piKL that regularizes a reward-maximizing policy toward a human imitationlearned policy. We prove that this is a no-regret learning algorithm under a modified utility function. We then show that DiL-piKL can be extended into a self-play reinforcement learning algorithm we call RL-DiL-piKL that provides a model of human play while simultaneously training an agent that responds well to this human model. We used RL-DiL-piKL to train an agent we name Diplodocus. In a 200-game no-press Diplomacy tournament involving 62 human participants spanning skill levels from beginner to expert, two Diplodocus agents both achieved a higher average score than all other participants who played more than two games, and ranked first and third according to an Elo ratings model. * Equal first author contribution. 1 Dota 2 is a two-team zero-sum game, but the presence of full information sharing between teammates makes it equivalent to 2p0s. Beyond 2p0s settings, self-play algorithms have also proven successful in highly adversarial games like six-player poker Brown & Sandholm (2019) .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- Language Instructed Reinforcement Learning for Human-AI CoordinationHengyuan Hu, Dorsa SadighICML 2023 · 被引用 90 次
- The Consensus Game: Language Model Generation via Equilibrium SearchAthul Paul Jacob, Yikang Shen, Gabriele Farina, Jacob AndreasICLR 2024 · 被引用 40 次
- Among Us: A Sandbox for Measuring and Detecting Agentic DeceptionSatvik Golechha, Adrià Garriga-AlonsoNeurIPS 2025 · 被引用 27 次
- Minimum Coverage Sets for Training Robust Ad Hoc Teamwork AgentsMuhammad Rahman, Jiaxun Cui, Peter StoneAAAI 2024 · 被引用 20 次
- Polynomial-Time Linear-Swap Regret Minimization in Imperfect-Information Sequential GamesGabriele Farina, Charilaos PipisNeurIPS 2023 · 被引用 13 次
它引用的顶会 Paper6
- "Other-Play" for Zero-Shot CoordinationHengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob N. FoersterICML 2020 · 被引用 271 次
- Off-Belief LearningHengyuan Hu, Adam Lerer, Brandon Cui, Luis Pineda 等ICML 2021 · 被引用 86 次
- Modeling Strong and Human-Like Gameplay with KL-Regularized SearchAthul Paul Jacob, David J. Wu, Gabriele Farina, Adam Lerer 等ICML 2022 · 被引用 69 次
- Human-Level Performance in No-Press Diplomacy via Equilibrium SearchJonathan Gray, Adam Lerer, Anton Bakhtin, Noam BrownICLR 2021 · 被引用 61 次
- No-Press Diplomacy from ScratchAnton Bakhtin, David J. Wu, Adam Lerer, Noam BrownNeurIPS 2021 · 被引用 51 次
相关 Paper
- Learning to Play No-Press Diplomacy with Best Response Policy IterationThomas W. Anthony, Tom Eccles, Andrea Tacchetti, János Kramár 等NeurIPS 2020 · 被引用 50 次
- DipLLM: Fine-Tuning LLM for Strategic Decision-making in DiplomacyKaixuan Xu, Jiajun Chai, Sicheng Li, Yuqian Fu 等ICML 2025
- Richelieu: Self-Evolving LLM-Based Agents for AI DiplomacyZhenyu Guan, Xiangyu Kong, Fangwei Zhong, Yizhou WangNeurIPS 2024 · 被引用 48 次
- Deep Reinforcement Learning for General Game PlayingAdrian Goldwaser, Michael ThielscherAAAI 2020 · 被引用 46 次
- Guarantees for Self-Play in Multiplayer Games via Polymatrix DecomposabilityRevan MacQueen, James R. WrightNeurIPS 2023 · 被引用 4 次
