Hierarchical Reinforcement Learning for Open-Domain Dialog
Abdelrhman Saleh, Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Rosalind W. Picard
Abstract
Open-domain dialog generation is a challenging problem; maximum likelihood training can lead to repetitive outputs, models have difficulty tracking long-term conversational goals, and training on standard movie or online datasets may lead to the generation of inappropriate, biased, or offensive text. Reinforcement Learning (RL) is a powerful framework that could potentially address these issues, for example by allowing a dialog model to optimize for reducing toxicity and repetitiveness. However, previous approaches which apply RL to open-domain dialog generation do so at the word level, making it difficult for the model to learn proper credit assignment for long-term conversational rewards. In this paper, we propose a novel approach to hierarchical reinforcement learning (HRL), VHRL, which uses policy gradients to tune the utterance-level embedding of a variational sequence model. This hierarchical approach provides greater flexibility for learning long-term, conversational rewards. We use self-play and RL to optimize for a set of human-centered conversation metrics, and show that our approach provides significant improvements – in terms of both human evaluation and automatic metrics – over state-of-the-art dialog models, including Transformers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9a16684b-d662-4617-bef3-d3b4e7604bb2Cited by top-tier papers15
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai et al.EMNLP 2022 · 239 citations
- Just Say No: Analyzing the Stance of Neural Dialogue Generation in Offensive ContextsAshutosh Baheti, Maarten Sap, Alan Ritter, Mark O. RiedlEMNLP 2021 · 50 citations
- Modelling Hierarchical Structure between Dialogue Policy and Natural Language Generator with Option Framework for Task-oriented Dialogue SystemJianhong Wang, Yuan Zhang, Tae-Kyun Kim, Yunjie GuICLR 2021 · 48 citations
- Dungeons and Dragons as a Dialog Challenge for Artificial IntelligenceChris Callison-Burch, Gaurav Singh Tomar, Lara J. Martin, Daphne Ippolito et al.EMNLP 2022 · 22 citations
- SelectAugment: Hierarchical Deterministic Sample Selection for Data AugmentationShiqi Lin, Zhizheng Zhang, Xin Li, Zhibo ChenAAAI 2023 · 13 citations
Related papers
- Knowledge Graph Grounded Goal Planning for Open-Domain Conversation GenerationJun Xu, Haifeng Wang, Zhengyu Niu, Hua Wu et al.AAAI 2020 · 69 citations
- Human-centric dialog training via offline reinforcement learningNatasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun, Craig Ferguson et al.EMNLP 2020 · 9 citations
- On the Effectiveness of Offline RL for Dialogue Response GenerationPaloma Sodhi, Felix Wu, Ethan R. Elenberg, Kilian Q. Weinberger et al.ICML 2023 · 6 citations
- Improving Multi-party Dialogue Generation via Topic and Rhetorical CoherenceYaxin Fan, Peifeng Li, Qiaoming ZhuEMNLP 2024 · 1 citation
- Building Persona Consistent Dialogue Agents with Offline Reinforcement LearningRyan Shea, Zhou YuEMNLP 2023 · 4 citations
