Semi-Supervised Dialogue Policy Learning via Stochastic Reward Estimation
Xinting Huang, Jianzhong Qi, Yu Sun, Rui Zhang
Abstract
Dialogue policy optimization often obtains feedback until task completion in taskoriented dialogue systems. This is insufficient for training intermediate dialogue turns since supervision signals (or rewards) are only provided at the end of dialogues. To address this issue, reward learning has been introduced to learn from state-action pairs of an optimal policy to provide turn-by-turn rewards. This approach requires complete state-action annotations of human-to-human dialogues (i.e., expert demonstrations), which is labor intensive. To overcome this limitation, we propose a novel reward learning approach for semisupervised policy learning. The proposed approach learns a dynamics model as the reward function which models dialogue progress (i.e., state-action sequences) based on expert demonstrations, either with or without annotations. The dynamics model computes rewards by predicting whether the dialogue progress is consistent with expert demonstrations. We further propose to learn action embeddings for a better generalization of the reward function. The proposed approach outperforms competitive policy learning baselines on MultiWOZ, a benchmark multi-domain dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09210d8e-09c4-4a0b-9c79-be83f37ff376Cited by top-tier papers3
- GraphDialog: Integrating Graph Knowledge into End-to-End Task-Oriented Dialogue SystemsShiquan Yang, Rui Zhang, Sarah M. ErfaniEMNLP 2020 · 46 citations
- An Interpretable Neuro-Symbolic Reasoning Framework for Task-Oriented Dialogue GenerationShiquan Yang, Rui Zhang, Sarah M. Erfani, Jey Han LauACL 2022 · 17 citations
- Multi-Action Dialog Policy Learning from Logged User FeedbackShuo Zhang, Junzhou Zhao, Pinghui Wang, Tianxiang Wang et al.AAAI 2023
Related papers
- GALAXY: A Generative Pre-trained Model for Task-Oriented Dialog with Semi-supervised Learning and Explicit Policy InjectionWanwei He, Yinpei Dai, Yinhe Zheng, Yuchuan Wu et al.AAAI 2022 · 181 citations
- [CASPI] Causal-aware Safe Policy Improvement for Task-oriented DialogueGovardana Sachithanandam Ramachandran, Kazuma Hashimoto, Caiming XiongACL 2022 · 12 citations
- A Probabilistic End-To-End Task-Oriented Dialog Model with Latent Belief States towards Semi-Supervised LearningYichi Zhang, Zhijian Ou, Min Hu, Junlan FengEMNLP 2020 · 52 citations
- KRLS: Improving End-to-End Response Generation in Task Oriented Dialog with Reinforced Keywords LearningXiao Yu, Qingyang Wu, Kun Qian, Zhou YuEMNLP 2023 · 1 citation
- Modelling Hierarchical Structure between Dialogue Policy and Natural Language Generator with Option Framework for Task-oriented Dialogue SystemJianhong Wang, Yuan Zhang, Tae-Kyun Kim, Yunjie GuICLR 2021 · 48 citations
