Modeling Strong and Human-Like Gameplay with KL-Regularized Search
Athul Paul Jacob, David J. Wu, Gabriele Farina, Adam Lerer, Hengyuan Hu, Anton Bakhtin, Jacob Andreas, Noam Brown
Abstract
We consider the task of building strong but human-like policies in multi-agent decision-making problems, given examples of human behavior. Imitation learning is effective at predicting human actions but may not match the strength of expert humans, while self-play learning and search techniques (e.g. AlphaZero) lead to strong performance but may produce policies that are difficult for humans to understand and coordinate with. We show in chess and Go that regularizing search based on the KL divergence from an imitation-learned policy results in higher human prediction accuracy and stronger performance than imitation learning alone. We then introduce a novel regret minimization algorithm that is regularized based on the KL divergence from an imitation-learned policy, and show that using this algorithm for search in no-press Diplomacy yields a policy that matches the human prediction accuracy of imitation learning while being substantially stronger.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers26
- Inverse Reinforcement Learning without Reinforcement LearningGokul Swamy, David Wu, Sanjiban Choudhury, Drew Bagnell et al.ICML 2023 · 49 citations
- Richelieu: Self-Evolving LLM-Based Agents for AI DiplomacyZhenyu Guan, Xiangyu Kong, Fangwei Zhong, Yizhou WangNeurIPS 2024 · 48 citations
- The Consensus Game: Language Model Generation via Equilibrium SearchAthul Paul Jacob, Yikang Shen, Gabriele Farina, Jacob AndreasICLR 2024 · 40 citations
- Maia-2: A Unified Model for Human-AI Alignment in ChessZhenwei Tang, Difan Jiao, Reid McIlroy-Young, Jon M. Kleinberg et al.NeurIPS 2024 · 39 citations
- Diverse Conventions for Human-AI CollaborationBidipta Sarkar, Andy Shih, Dorsa SadighNeurIPS 2023 · 23 citations
Builds on13
- Keep Doing What Worked: Behavior Modelling Priors for Offline Reinforcement LearningNoah Y. Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki et al.ICLR 2020 · 299 citations
- "Other-Play" for Zero-Shot CoordinationHengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob N. FoersterICML 2020 · 271 citations
- Fast Policy Extragradient Methods for Competitive Games with Entropy RegularizationShicong Cen, Yuting Wei, Yuejie ChiNeurIPS 2021 · 105 citations
- Improving Policies via Search in Cooperative Partially Observable GamesAdam Lerer, Hengyuan Hu, Jakob N. Foerster, Noam BrownAAAI 2020 · 87 citations
- Off-Belief LearningHengyuan Hu, Adam Lerer, Brandon Cui, Luis Pineda et al.ICML 2021 · 86 citations
Related papers
- Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and PlanningAnton Bakhtin, David J. Wu, Adam Lerer, Jonathan Gray et al.ICLR 2023 · 10 citations
- Human-Level Performance in No-Press Diplomacy via Equilibrium SearchJonathan Gray, Adam Lerer, Anton Bakhtin, Noam BrownICLR 2021 · 61 citations
- Policy improvement by planning with GumbelIvo Danihelka, Arthur Guez, Julian Schrittwieser, David SilverICLR 2022 · 84 citations
- Regret-Guided Search Control for Efficient Learning in AlphaZeroYun-Jui Tsai, Wei-Yu Chen, Yan-Ru Ju, Yu-Hung Chang et al.ICLR 2026
- Monte-Carlo Tree Search as Regularized Policy OptimizationJean-Bastien Grill, Florent Altché, Yunhao Tang, Thomas Hubert et al.ICML 2020 · 84 citations
