Taming Continuous Posteriors for Latent Variational Dialogue Policies
Marin Vlastelica, Patrick Ernst, Gyuri Szarvas
Abstract
Utilizing amortized variational inference for latent-action reinforcement learning (RL) has been shown to be an effective approach in Task-oriented Dialogue (ToD) systems for optimizing dialogue success.Until now, categorical posteriors have been argued to be one of the main drivers of performance. In this work we revisit Gaussian variational posteriors for latent-action RL and show that they can yield even better performance than categoricals. We achieve this by introducing an improved variational inference objective for learning continuous representations without auxiliary learning objectives, which streamlines the training procedure. Moreover, we propose ways to regularize the latent dialogue policy, which helps to retain good response coherence. Using continuous latent representations our model achieves state of the art dialogue success rate on the MultiWOZ benchmark, and also compares well to categorical latent methods in response coherence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2de4af5-9303-4fca-bcbb-93abff83b8a5Builds on6
- Task-Oriented Dialog Systems That Consider Multiple Appropriate Responses under the Same ContextYichi Zhang, Zhijian Ou, Zhou YuAAAI 2020 · 198 citations
- CoCo: Controllable Counterfactuals for Evaluating Dialogue State TrackersShiyang Li, Semih Yavuz, Kazuma Hashimoto, Jia Li et al.ICLR 2021 · 65 citations
- Hierarchical Reinforcement Learning for Open-Domain DialogAbdelrhman Saleh, Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen et al.AAAI 2020 · 60 citations
- Multi-Domain Dialogue Acts and Response Co-GenerationKai Wang, Junfeng Tian, Rui Wang, Xiaojun Quan et al.ACL 2020 · 46 citations
- UniConv: A Unified Conversational Neural Architecture for Multi-domain Task-oriented DialoguesHung Le, Doyen Sahoo, Chenghao Liu, Nancy F. Chen et al.EMNLP 2020 · 31 citations
Related papers
- Modelling Hierarchical Structure between Dialogue Policy and Natural Language Generator with Option Framework for Task-oriented Dialogue SystemJianhong Wang, Yuan Zhang, Tae-Kyun Kim, Yunjie GuICLR 2021 · 48 citations
- KRLS: Improving End-to-End Response Generation in Task Oriented Dialog with Reinforced Keywords LearningXiao Yu, Qingyang Wu, Kun Qian, Zhou YuEMNLP 2023 · 1 citation
- MALA: Cross-Domain Dialogue Generation with Action LearningXinting Huang, Jianzhong Qi, Yu Sun, Rui ZhangAAAI 2020 · 19 citations
- TA&AT: Enhancing Task-Oriented Dialog with Turn-Level Auxiliary Tasks and Action-Tree Based Scheduled SamplingLongxiang Liu, Xiuxing Li, Yang FengAAAI 2024 · 1 citation
- Iterative Amortized Policy OptimizationJoseph Marino, Alexandre Piché, Alessandro Davide Ialongo, Yisong YueNeurIPS 2021 · 27 citations
