Large Scale Multi-Actor Generative Dialog Modeling
Alex Boyd, Raul Puri, Mohammad Shoeybi, Mostofa Patwary, Bryan Catanzaro
Abstract
Non-goal oriented dialog agents (i.e. chatbots) aim to produce varying and engaging conversations with a user; however, they typically exhibit either inconsistent personality across conversations or the average personality of all users. This paper addresses these issues by controlling an agent's persona upon generation via conditioning on prior conversations of a target actor. In doing so, we are able to utilize more abstract patterns within a person's speech and better emulate them in generated responses. This work introduces the GENERATIVE CONVERSATION CONTROL model, an augmented and fine-tuned GPT-2 language model that conditions on past reference conversations to probabilistically model multi-turn conversations in the actor's persona. We introduce an accompanying data collection procedure to obtain 10.3M conversations from 6 months worth of Reddit comments. We demonstrate that scaling model sizes from 117M to 8.3B parameters yields an improvement from 23.14 to 13.14 perplexity on 1.7M held out Reddit conversations. Increasing model scale yielded similar improvements in human evaluations that measure preference of model samples to the held out target distribution in terms of realism (31% increased to 37% preference), style matching (37% to 42%), grammar and content quality (29% to 42%), and conversation coherency (32% to 40%). We find that conditionally modeling past conversations improves perplexity by 0.47 in automatic evaluations. Through human trials we identify positive trends between conditional modeling and style matching and outline steps to further improve persona control. * First two authors have contributed equally. † Research conducted during an internship at NVIDIA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- MEGATRON-CNTRL: Controllable Story Generation with External Knowledge Using Large-Scale Language ModelsPeng Xu, Mostofa Patwary, Mohammad Shoeybi, Raul Puri et al.EMNLP 2020 · 104 citations
- Call for Customized Conversation: Customized Conversation Grounding Persona and KnowledgeYoonna Jang, Jungwoo Lim, Yuna Hur, Dongsuk Oh et al.AAAI 2022 · 47 citations
- Training Question Answering Models From Synthetic DataRaul Puri, Ryan Spring, Mohammad Shoeybi, Mostofa Patwary et al.EMNLP 2020 · 15 citations
- SimOAP: Improve Coherence and Consistency in Persona-based Dialogue Generation via Over-sampling and Post-evaluationJunkai Zhou, Liang Pang, Huawei Shen, Xueqi ChengACL 2023 · 6 citations
- X-TURING: Towards an Enhanced and Efficient Turing Test for Long-Term Dialogue AgentsWeiqi Wu, Hongqiu Wu, Hai ZhaoACL 2025 · 6 citations
Related papers
- The Personality Dimensions GPT-3 Expresses During Human-Chatbot InteractionsNikola Kovacevic, Christian Holz, Markus Gross, Rafael WampflerUbiComp 2024 · 15 citations
- Dialogue Response Ranking Training with Large-Scale Human Feedback DataXiang Gao, Yizhe Zhang, Michel Galley, Chris Brockett et al.EMNLP 2020 · 67 citations
- A Disentangled-Attention Based Framework with Persona-Aware Prompt Learning for Dialogue GenerationPingsheng Liu, Zhengjie Huang, Xiechi Zhang, Linlin Wang et al.AAAI 2023 · 8 citations
- BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded DataWenkai Li, Jiarui Liu, Andy Liu, Xuhui Zhou et al.ACL 2025
- Generative Expressive Conversational Speech SynthesisRui Liu, Yifan Hu, Yi Ren, Xiang Yin et al.ACM MM 2024 · 15 citations
