DipLLM: Fine-Tuning LLM for Strategic Decision-making in Diplomacy
Kaixuan Xu, Jiajun Chai, Sicheng Li, Yuqian Fu, Yuanheng Zhu, Dongbin Zhao
摘要
Diplomacy is a complex multiplayer game that requires both cooperation and competition, posing significant challenges for AI systems. Traditional methods rely on equilibrium search to generate extensive game data for training, which demands substantial computational resources. Large Language Models (LLMs) offer a promising alternative, leveraging pre-trained knowledge to achieve strong performance with relatively small-scale fine-tuning. However, applying LLMs to Diplomacy remains challenging due to the exponential growth of possible action combinations and the intricate strategic interactions among players. To address this challenge, we propose DipLLM, a fine-tuned LLM-based agent that learns equilibrium policies for Diplomacy. DipLLM employs an autoregressive factorization framework to simplify the complex task of multi-unit action assignment into a sequence of unit-level decisions. By defining an equilibrium policy within this framework as the learning objective, we fine-tune the model using only 1.5% of the data required by the state-of-the-art Cicero model, surpassing its performance. Our results demonstrate the potential of fine-tuned LLMs for tackling complex strategic decision-making in multiplayer games.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Empowering Multi-Robot Cooperation via Sequential World ModelsZijie Zhao, Honglei Guo, Shengqian Chen, Kaixuan Xu 等ICLR 2026 · 被引用 16 次
- Scaling Inference-Time Computation via Opponent Simulation: Enabling Online Strategic Adaptation in Repeated NegotiationXiangyu Liu, Di Wang, Zhe Feng, Aranyak MehtaICML 2026 · 被引用 2 次
- Generative Gamer: Learning Equilibrium Strategy by LLM-driven Dynamic DeductionYadong Zhang, Xinshu Shen, Yupei Ren, Shangqing Zhao 等ACL 2026
- Multi-agent KTO: Enhancing Strategic Interactions of Large Language Model in Language GameRong Ye, Yongxin Zhang, Yikai Zhang, Haoyu Kuang 等NeurIPS 2025
它引用的顶会 Paper15
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement LearningHao Bai, Yifei Zhou, Jiayi Pan, Mert Cemri 等NeurIPS 2024 · 被引用 239 次
- Language Agents with Reinforcement Learning for Strategic Play in the Werewolf GameZelai Xu, Chao Yu, Fei Fang, Yu Wang 等ICML 2024 · 被引用 145 次
- Modeling Strong and Human-Like Gameplay with KL-Regularized SearchAthul Paul Jacob, David J. Wu, Gabriele Farina, Adam Lerer 等ICML 2022 · 被引用 69 次
相关 Paper
- Richelieu: Self-Evolving LLM-Based Agents for AI DiplomacyZhenyu Guan, Xiangyu Kong, Fangwei Zhong, Yizhou WangNeurIPS 2024 · 被引用 48 次
- Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and PlanningAnton Bakhtin, David J. Wu, Adam Lerer, Jonathan Gray 等ICLR 2023 · 被引用 10 次
- Empowering Economic Simulation for Massively Multiplayer Online Games through Generative Agent-Based ModelingBihan Xu, Shiwei Zhao, Runze Wu, Zhenya Huang 等KDD 2025 · 被引用 2 次
- Learning to Play No-Press Diplomacy with Best Response Policy IterationThomas W. Anthony, Tom Eccles, Andrea Tacchetti, János Kramár 等NeurIPS 2020 · 被引用 50 次
- Systematic Biases in LLM Simulations of DebatesAmir Taubenfeld, Yaniv Dover, Roi Reichart, Ariel GoldsteinEMNLP 2024 · 被引用 35 次
