Efficient Dialog Policy Learning by Reasoning with Contextual Knowledge
Haodi Zhang, Zhichao Zeng, Keting Lu, Kaishun Wu, Shiqi Zhang
Abstract
Goal-oriented dialog policy learning algorithms aim to learn a dialog policy for selecting language actions based on the current dialog state. Deep reinforcement learning methods have been used for dialog policy learning. This work is motivated by the observation that, although dialog is a domain with rich contextual knowledge, reinforcement learning methods are ill-equipped to incorporate such knowledge into the dialog policy learning process. In this paper, we develop a deep reinforcement learning framework for goal-oriented dialog policy learning that learns user preferences from user goal data, while leveraging commonsense knowledge from people. The developed framework has been evaluated using a realistic dialog simulation platform. Compared with baselines from the literature and the ablations of our approach, we see significant improvements in learning efficiency and the quality of the computed action policies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- Multi-Agent Task-Oriented Dialog Policy Learning with Role-Aware Reward DecompositionRyuichi Takanobu, Runze Liang, Minlie HuangACL 2020 · 47 citations
- Automatic Curriculum Learning With Over-repetition Penalty for Dialogue Policy LearningYangyang Zhao, Zhenyu Wang, Zhenhua HuangAAAI 2021 · 20 citations
- [CASPI] Causal-aware Safe Policy Improvement for Task-oriented DialogueGovardana Sachithanandam Ramachandran, Kazuma Hashimoto, Caiming XiongACL 2022 · 12 citations
- Human-centric dialog training via offline reinforcement learningNatasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun, Craig Ferguson et al.EMNLP 2020 · 9 citations
- Knowledge Graph Grounded Goal Planning for Open-Domain Conversation GenerationJun Xu, Haifeng Wang, Zhengyu Niu, Hua Wu et al.AAAI 2020 · 69 citations
