Building Persona Consistent Dialogue Agents with Offline Reinforcement Learning
Ryan Shea, Zhou Yu
Abstract
Maintaining a consistent persona is a key quality for any open domain dialogue system. Current state-of-the-art systems do this by training agents with supervised learning or online reinforcement learning (RL). However, systems trained with supervised learning often lack consistency as they are never punished for uttering contradictions. Additional training with RL can alleviate some of these issues, however the training process is expensive. Instead, we propose an offline RL framework to improve the persona consistency of dialogue systems. Our framework allows us to combine the advantages of previous methods as we can inexpensively train our model on existing data as in supervised learning, while punishing and rewarding specific utterances as in RL. We also introduce a simple importance sampling method to reduce the variance of importance weights in offline RL training which we call Variance-Reducing MLE-Initialized (VaRMI) importance sampling. Our automatic and human evaluations show that our framework improves both the persona consistency and dialogue quality of a state-of-the-art social chatbot.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c15c604a-bade-49c5-b900-fb25495f7f5cCited by top-tier papers3
- Consistently Simulating Human Personas with Multi-Turn Reinforcement LearningMarwa Abdulhai, Ryan Cheng, Donovan Clay, Tim Althoff et al.NeurIPS 2025 · 51 citations
- RL-GPT: Integrating Reinforcement Learning and Code-as-policyShaoteng Liu, Haoqi Yuan, Minda Hu, Yanwei Li et al.NeurIPS 2024 · 48 citations
- DialogLab: Authoring, Simulating, and Testing Dynamic Human-AI Group ConversationsErzhen Hu, Yanhe Chen, Mingyi Li, Vrushank Phadnis et al.UIST 2025 · 3 citations
Builds on10
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
- You Impress Me: Dialogue Generation via Mutual Persona PerceptionQian Liu, Yihong Chen, Bei Chen, Jian-Guang Lou et al.ACL 2020 · 144 citations
- Don't Say That! Making Inconsistent Dialogue Unlikely with Unlikelihood TrainingMargaret Li, Stephen Roller, Ilia Kulikov, Sean Welleck et al.ACL 2020 · 120 citations
- Text Generation by Learning from DemonstrationsRichard Yuanzhe Pang, He HeICLR 2021 · 88 citations
- Generate, Delete and Rewrite: A Three-Stage Framework for Improving Persona Consistency of Dialogue GenerationHaoyu Song, Yan Wang, Weinan Zhang, Xiaojiang Liu et al.ACL 2020 · 86 citations
Related papers
- Learning to Know Myself: A Coarse-to-Fine Persona-Aware Training Framework for Personalized Dialogue GenerationYunpeng Li, Yue Hu, Yajing Sun, Luxi Xing et al.AAAI 2023 · 9 citations
- Human-centric dialog training via offline reinforcement learningNatasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun, Craig Ferguson et al.EMNLP 2020 · 9 citations
- Hierarchical Reinforcement Learning for Open-Domain DialogAbdelrhman Saleh, Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen et al.AAAI 2020 · 60 citations
- On the Effectiveness of Offline RL for Dialogue Response GenerationPaloma Sodhi, Felix Wu, Ethan R. Elenberg, Kilian Q. Weinberger et al.ICML 2023 · 6 citations
- Will I Sound Like Me? Improving Persona Consistency in Dialogues through Pragmatic Self-ConsciousnessHyunwoo Kim, Byeongchang Kim, Gunhee KimEMNLP 2020 · 46 citations
