Dungeons and Dragons as a Dialog Challenge for Artificial Intelligence
Chris Callison-Burch, Gaurav Singh Tomar, Lara J. Martin, Daphne Ippolito, Suma Bailis, David Reitter
Abstract
AI researchers have posited Dungeons and Dragons (D&D) as a challenge problem to test systems on various language-related capabilities. In this paper, we frame D&D specifically as a dialogue system challenge, where the tasks are to both generate the next conversational turn in the game and predict the state of the game given the dialogue history. We create a gameplay dataset consisting of nearly 900 games, with a total of 7,000 players, 800,000 dialogue turns, 500,000 dice rolls, and 58 million words. We automatically annotate the data with partial state information about the game play. We train a large language model (LM) to generate the next game turn, conditioning it on different information. The LM can respond as a particular character or as the player who runs the game—i.e., the Dungeon Master (DM). It is trained to produce dialogue that is either in-character (roleplaying in the fictional world) or out-of-character (discussing rules or strategy). We perform a human evaluation to determine what factors make the generated output plausible and interesting. We further perform an automatic evaluation to determine how well the model can predict the game state given the history and examine how well tracking the game state improves its ability to produce plausible conversational output.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b31de6f1-bd03-4e57-b773-302e2038f06eCited by top-tier papers7
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Character-LLM: A Trainable Agent for Role-PlayingYunfan Shao, Linyang Li, Junqi Dai, Xipeng QiuEMNLP 2023 · 97 citations
- I Cast Detect Thoughts: Learning to Converse and Guide with Intents and Theory-of-Mind in Dungeons and DragonsPei Zhou, Andrew Zhu, Jennifer Hu, Jay Pujara et al.ACL 2023 · 8 citations
- SimSpark: Interactive Simulation of Social Media BehaviorsZiyue Lin, Yi Shan, Lin Gao, Xinghua Jia et al.CSCW 2025 · 5 citations
- Ontologically Faithful Generation of Non-Player Character DialoguesNathaniel Weir, Ryan Thomas, Randolph D'Amore, Kellie Hill et al.EMNLP 2024 · 2 citations
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- PlotMachines: Outline-Conditioned Generation with Dynamic Plot State TrackingHannah Rashkin, Asli Celikyilmaz, Yejin Choi, Jianfeng GaoEMNLP 2020 · 100 citations
- Keep CALM and Explore: Language Models for Action Generation in Text-based GamesShunyu Yao, Rohan Rao, Matthew J. Hausknecht, Karthik NarasimhanEMNLP 2020 · 67 citations
- Hierarchical Reinforcement Learning for Open-Domain DialogAbdelrhman Saleh, Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen et al.AAAI 2020 · 60 citations
- Storytelling with Dialogue: A Critical Role Dungeons and Dragons DatasetRevanth Rameshkumar, Peter BaileyACL 2020 · 37 citations
Related papers
- FIREBALL: A Dataset of Dungeons and Dragons Actual-Play with Structured Game State InformationAndrew Zhu, Karmanya Aggarwal, Alexander H. Feng, Lara J. Martin et al.ACL 2023 · 9 citations
- Personalized Quest and Dialogue Generation in Role-Playing Games: A Knowledge Graph- and Language Model-based ApproachTrevor Ashby, Braden K. Webb, Gregory Knapp, Jackson Searle et al.CHI 2023 · 64 citations
- DMT-RoleBench: A Dynamic Multi-Turn Dialogue Based Benchmark for Role-Playing Evaluation of Large Language Model and AgentDingbo Yuan, Yipeng Chen, Guodong Liu, Chenchen Li et al.AAAI 2025 · 6 citations
- Don't Stop the Multi-Party! On Generating Synthetic Written Multi-Party Conversations with ConstraintsNicolò Penzo, Marco Guerini, Bruno Lepri, Goran Glavas et al.AAAI 2026 · 3 citations
- LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language ModelsMarwa Abdulhai, Isadora White, Charlie Victor Snell, Charles Sun et al.ICML 2025
