Hierarchical Reinforcement Learning for Open-Domain Dialog
Abdelrhman Saleh, Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Rosalind W. Picard
摘要
Open-domain dialog generation is a challenging problem; maximum likelihood training can lead to repetitive outputs, models have difficulty tracking long-term conversational goals, and training on standard movie or online datasets may lead to the generation of inappropriate, biased, or offensive text. Reinforcement Learning (RL) is a powerful framework that could potentially address these issues, for example by allowing a dialog model to optimize for reducing toxicity and repetitiveness. However, previous approaches which apply RL to open-domain dialog generation do so at the word level, making it difficult for the model to learn proper credit assignment for long-term conversational rewards. In this paper, we propose a novel approach to hierarchical reinforcement learning (HRL), VHRL, which uses policy gradients to tune the utterance-level embedding of a variational sequence model. This hierarchical approach provides greater flexibility for learning long-term, conversational rewards. We use self-play and RL to optimize for a set of human-centered conversation metrics, and show that our approach provides significant improvements – in terms of both human evaluation and automatic metrics – over state-of-the-art dialog models, including Transformers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai 等EMNLP 2022 · 被引用 239 次
- Just Say No: Analyzing the Stance of Neural Dialogue Generation in Offensive ContextsAshutosh Baheti, Maarten Sap, Alan Ritter, Mark O. RiedlEMNLP 2021 · 被引用 50 次
- Modelling Hierarchical Structure between Dialogue Policy and Natural Language Generator with Option Framework for Task-oriented Dialogue SystemJianhong Wang, Yuan Zhang, Tae-Kyun Kim, Yunjie GuICLR 2021 · 被引用 48 次
- Dungeons and Dragons as a Dialog Challenge for Artificial IntelligenceChris Callison-Burch, Gaurav Singh Tomar, Lara J. Martin, Daphne Ippolito 等EMNLP 2022 · 被引用 22 次
- SelectAugment: Hierarchical Deterministic Sample Selection for Data AugmentationShiqi Lin, Zhizheng Zhang, Xin Li, Zhibo ChenAAAI 2023 · 被引用 13 次
相关 Paper
- Knowledge Graph Grounded Goal Planning for Open-Domain Conversation GenerationJun Xu, Haifeng Wang, Zhengyu Niu, Hua Wu 等AAAI 2020 · 被引用 69 次
- Human-centric dialog training via offline reinforcement learningNatasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun, Craig Ferguson 等EMNLP 2020 · 被引用 9 次
- On the Effectiveness of Offline RL for Dialogue Response GenerationPaloma Sodhi, Felix Wu, Ethan R. Elenberg, Kilian Q. Weinberger 等ICML 2023 · 被引用 6 次
- Improving Multi-party Dialogue Generation via Topic and Rhetorical CoherenceYaxin Fan, Peifeng Li, Qiaoming ZhuEMNLP 2024 · 被引用 1 次
- Building Persona Consistent Dialogue Agents with Offline Reinforcement LearningRyan Shea, Zhou YuEMNLP 2023 · 被引用 4 次
