More Victories, Less Cooperation: Assessing Cicero's Diplomacy Play
Wichayaporn Wongkamjan, Feng Gu, Yanze Wang, Ulf Hermjakob, Jonathan May, Brandon M. Stewart, Jonathan K. Kummerfeld, Denis Peskoff, Jordan L. Boyd-Graber
Abstract
The boardgame Diplomacy is a challenging setting for communicative and cooperative artificial intelligence. The most prominent communicative Diplomacy AI, Cicero, has excellent strategic abilities, exceeding human players. However, the best Diplomacy players master communication, not just tactics, which is why the game has received attention as an AI challenge. This work seeks to understand the degree to which Cicero succeeds at communication. First, we annotate in-game communication with abstract meaning representation to separate in-game tactics from general language. Second, we run two dozen games with humans and Cicero, totaling over 200 humanplayer hours of competition. While AI can consistently outplay human players, AI-Human communication is still limited because of AI's difficulty with deception and persuasion. This shows that Cicero relies on strategy and has not yet reached the full promise of communicative and cooperative AI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ad8078e2-2ec4-41c6-b7ad-5a4747966dffCited by top-tier papers5
- CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social DilemmasEmanuel Tewolde, Xiao Zhang, David Guzman Piedrahita, Vincent Conitzer et al.ICML 2026 · 15 citations
- Explaining Decisions of Agents in Mixed-Motive GamesMaayan Orner, Oleg Maksimov, Akiva Kleinerman, Charles Ortiz et al.AAAI 2025 · 4 citations
- NewsInterview: a Dataset and a Playground to Evaluate LLMs' Grounding Gap via Informational InterviewsAlexander Spangher, Michael Lu, Sriya Kalyan, Hyundong Justin Cho et al.ACL 2025 · 2 citations
- DipLLM: Fine-Tuning LLM for Strategic Decision-making in DiplomacyKaixuan Xu, Jiajun Chai, Sicheng Li, Yuqian Fu et al.ICML 2025
- Mental Model-based Generation of Lies for Insider Threat ModelingBrittany Cates, Sarath SreedharanAAAI 2026
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model PromptsTongshuang Wu, Michael Terry, Carrie Jun CaiCHI 2022 · 465 citations
- Formalizing and Benchmarking Prompt Injection Attacks and DefensesYupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia et al.USENIX Security 2024 · 308 citations
Related papers
- Human-Level Performance in No-Press Diplomacy via Equilibrium SearchJonathan Gray, Adam Lerer, Anton Bakhtin, Noam BrownICLR 2021 · 61 citations
- Learning to Play No-Press Diplomacy with Best Response Policy IterationThomas W. Anthony, Tom Eccles, Andrea Tacchetti, János Kramár et al.NeurIPS 2020 · 50 citations
- It Takes Two to Negotiate: Modeling Social Exchange in Online Multiplayer GamesKokil Jaidka, Hansin Ahuja, Lynnette Hui Xian NgCSCW 2024 · 9 citations
- The Hidden Rules of Hanabi: How Humans Outperform AI AgentsMatthew Sidji, Wally Smith, Melissa J. RogersonCHI 2023 · 9 citations
- It Takes Two to Lie: One to Lie, and One to ListenDenis Peskov, Benny Cheng, Ahmed Elgohary, Joe Barrow et al.ACL 2020 · 28 citations
