Efficient Dialogue Complementary Policy Learning via Deep Q-network Policy and Episodic Memory Policy
Yangyang Zhao, Zhenyu Wang, Changxi Zhu, Shihan Wang
Abstract
Deep reinforcement learning has shown great potential in training dialogue policies. However, its favorable performance comes at the cost of many rounds of interaction. Most of the existing dialogue policy methods rely on a single learning system, while the human brain has two specialized learning and memory systems, supporting to find good solutions without requiring copious examples. Inspired by the human brain, this paper proposes a novel complementary policy learning (CPL) framework, which exploits the complementary advantages of the episodic memory (EM) policy and the deep Q-network (DQN) policy to achieve fast and effective dialogue policy learning. In order to coordinate between the two policies, we proposed a confidence controller to control the complementary time according to their relative efficacy at different stages. Furthermore, memory connectivity and time pruning are proposed to guarantee the flexible and adaptive generalization of the EM policy in dialog tasks. Experimental results on three dialogue datasets show that our method significantly outperforms existing methods relying on a single learning system.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a5f51655-8575-4338-83ee-162d8a3deddaCited by top-tier papers1
Ask how each one uses itBuilds on3
- Dynamic Reward-Based Dueling Deep Dyna-Q: Robust Policy Learning in Noisy EnvironmentsYangyang Zhao, Zhenyu Wang, Kai Yin, Rui Zhang et al.AAAI 2020 · 32 citations
- Automatic Curriculum Learning With Over-repetition Penalty for Dialogue Policy LearningYangyang Zhao, Zhenyu Wang, Zhenhua HuangAAAI 2021 · 20 citations
- Learning Efficient Dialogue Policy from Demonstrations through ShapingHuimin Wang, Baolin Peng, Kam-Fai WongACL 2020 · 18 citations
Related papers
- Multi-Agent Task-Oriented Dialog Policy Learning with Role-Aware Reward DecompositionRyuichi Takanobu, Runze Liang, Minlie HuangACL 2020 · 47 citations
- Learning Multi-Agent Communication through Structured Attentive ReasoningMurtaza Rangwala, Ryan WilliamsNeurIPS 2020 · 42 citations
- Task-Completion Dialogue Policy Learning via Monte Carlo Tree Search with Dueling NetworkSihan Wang, Kaijie Zhou, Kunfeng Lai, Jianping ShenEMNLP 2020 · 10 citations
- Episodic Multi-agent Reinforcement Learning with Curiosity-driven ExplorationLulu Zheng, Jiarui Chen, Jianhao Wang, Jiamin He et al.NeurIPS 2021 · 126 citations
- Learning Fast, Learning Slow: A General Continual Learning Method based on Complementary Learning SystemElahe Arani, Fahad Sarfraz, Bahram ZonoozICLR 2022 · 168 citations
