Building Persona Consistent Dialogue Agents with Offline Reinforcement Learning
Ryan Shea, Zhou Yu
摘要
Maintaining a consistent persona is a key quality for any open domain dialogue system. Current state-of-the-art systems do this by training agents with supervised learning or online reinforcement learning (RL). However, systems trained with supervised learning often lack consistency as they are never punished for uttering contradictions. Additional training with RL can alleviate some of these issues, however the training process is expensive. Instead, we propose an offline RL framework to improve the persona consistency of dialogue systems. Our framework allows us to combine the advantages of previous methods as we can inexpensively train our model on existing data as in supervised learning, while punishing and rewarding specific utterances as in RL. We also introduce a simple importance sampling method to reduce the variance of importance weights in offline RL training which we call Variance-Reducing MLE-Initialized (VaRMI) importance sampling. Our automatic and human evaluations show that our framework improves both the persona consistency and dialogue quality of a state-of-the-art social chatbot.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Consistently Simulating Human Personas with Multi-Turn Reinforcement LearningMarwa Abdulhai, Ryan Cheng, Donovan Clay, Tim Althoff 等NeurIPS 2025 · 被引用 51 次
- RL-GPT: Integrating Reinforcement Learning and Code-as-policyShaoteng Liu, Haoqi Yuan, Minda Hu, Yanwei Li 等NeurIPS 2024 · 被引用 48 次
- DialogLab: Authoring, Simulating, and Testing Dynamic Human-AI Group ConversationsErzhen Hu, Yanhe Chen, Mingyi Li, Vrushank Phadnis 等UIST 2025 · 被引用 3 次
它引用的顶会 Paper10
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan 等ICLR 2020 · 被引用 683 次
- You Impress Me: Dialogue Generation via Mutual Persona PerceptionQian Liu, Yihong Chen, Bei Chen, Jian-Guang Lou 等ACL 2020 · 被引用 144 次
- Don't Say That! Making Inconsistent Dialogue Unlikely with Unlikelihood TrainingMargaret Li, Stephen Roller, Ilia Kulikov, Sean Welleck 等ACL 2020 · 被引用 120 次
- Text Generation by Learning from DemonstrationsRichard Yuanzhe Pang, He HeICLR 2021 · 被引用 88 次
- Generate, Delete and Rewrite: A Three-Stage Framework for Improving Persona Consistency of Dialogue GenerationHaoyu Song, Yan Wang, Weinan Zhang, Xiaojiang Liu 等ACL 2020 · 被引用 86 次
相关 Paper
- Learning to Know Myself: A Coarse-to-Fine Persona-Aware Training Framework for Personalized Dialogue GenerationYunpeng Li, Yue Hu, Yajing Sun, Luxi Xing 等AAAI 2023 · 被引用 9 次
- Human-centric dialog training via offline reinforcement learningNatasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun, Craig Ferguson 等EMNLP 2020 · 被引用 9 次
- Hierarchical Reinforcement Learning for Open-Domain DialogAbdelrhman Saleh, Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen 等AAAI 2020 · 被引用 60 次
- On the Effectiveness of Offline RL for Dialogue Response GenerationPaloma Sodhi, Felix Wu, Ethan R. Elenberg, Kilian Q. Weinberger 等ICML 2023 · 被引用 6 次
- Will I Sound Like Me? Improving Persona Consistency in Dialogues through Pragmatic Self-ConsciousnessHyunwoo Kim, Byeongchang Kim, Gunhee KimEMNLP 2020 · 被引用 46 次
