Taming Continuous Posteriors for Latent Variational Dialogue Policies
Marin Vlastelica, Patrick Ernst, Gyuri Szarvas
摘要
Utilizing amortized variational inference for latent-action reinforcement learning (RL) has been shown to be an effective approach in Task-oriented Dialogue (ToD) systems for optimizing dialogue success.Until now, categorical posteriors have been argued to be one of the main drivers of performance. In this work we revisit Gaussian variational posteriors for latent-action RL and show that they can yield even better performance than categoricals. We achieve this by introducing an improved variational inference objective for learning continuous representations without auxiliary learning objectives, which streamlines the training procedure. Moreover, we propose ways to regularize the latent dialogue policy, which helps to retain good response coherence. Using continuous latent representations our model achieves state of the art dialogue success rate on the MultiWOZ benchmark, and also compares well to categorical latent methods in response coherence.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Task-Oriented Dialog Systems That Consider Multiple Appropriate Responses under the Same ContextYichi Zhang, Zhijian Ou, Zhou YuAAAI 2020 · 被引用 198 次
- CoCo: Controllable Counterfactuals for Evaluating Dialogue State TrackersShiyang Li, Semih Yavuz, Kazuma Hashimoto, Jia Li 等ICLR 2021 · 被引用 65 次
- Hierarchical Reinforcement Learning for Open-Domain DialogAbdelrhman Saleh, Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen 等AAAI 2020 · 被引用 60 次
- Multi-Domain Dialogue Acts and Response Co-GenerationKai Wang, Junfeng Tian, Rui Wang, Xiaojun Quan 等ACL 2020 · 被引用 46 次
- UniConv: A Unified Conversational Neural Architecture for Multi-domain Task-oriented DialoguesHung Le, Doyen Sahoo, Chenghao Liu, Nancy F. Chen 等EMNLP 2020 · 被引用 31 次
相关 Paper
- Modelling Hierarchical Structure between Dialogue Policy and Natural Language Generator with Option Framework for Task-oriented Dialogue SystemJianhong Wang, Yuan Zhang, Tae-Kyun Kim, Yunjie GuICLR 2021 · 被引用 48 次
- KRLS: Improving End-to-End Response Generation in Task Oriented Dialog with Reinforced Keywords LearningXiao Yu, Qingyang Wu, Kun Qian, Zhou YuEMNLP 2023 · 被引用 1 次
- MALA: Cross-Domain Dialogue Generation with Action LearningXinting Huang, Jianzhong Qi, Yu Sun, Rui ZhangAAAI 2020 · 被引用 19 次
- TA&AT: Enhancing Task-Oriented Dialog with Turn-Level Auxiliary Tasks and Action-Tree Based Scheduled SamplingLongxiang Liu, Xiuxing Li, Yang FengAAAI 2024 · 被引用 1 次
- Iterative Amortized Policy OptimizationJoseph Marino, Alexandre Piché, Alessandro Davide Ialongo, Yisong YueNeurIPS 2021 · 被引用 27 次
