Dynamic Reward-Based Dueling Deep Dyna-Q: Robust Policy Learning in Noisy Environments
Yangyang Zhao, Zhenyu Wang, Kai Yin, Rui Zhang, Zhenhua Huang, Pei Wang
Abstract
Task-oriented dialogue systems provide a convenient interface to help users complete tasks. An important consideration for task-oriented dialogue systems is the ability to against the noise commonly existed in the real-world conversation. Both rule-based strategies and statistical modeling techniques can solve noise problems, but they are costly. In this paper, we propose a new approach, called Dynamic Reward-based Dueling Deep Dyna-Q (DR-D3Q). The DR-D3Q can learn policies in noise robustly, and it is easy to implement by combining dynamic reward and the Dueling Deep Q-Network (Dueling DQN) into Deep Dyna-Q (DDQ) framework. The Dueling DQN can mitigate the negative impact of noise on learning policies, but it is inapplicable to dialogue domain due to different reward mechanisms. Unlike typical dialogue reward function, we integrate dynamic reward that provides reward in real-time for agent to make Dueling DQN adapt to dialogue domain. For the purpose of supplementing the limited amount of real user experiences, we take the DDQ framework as the basic framework. Experiments using simulation and human evaluation show that the DR-D3Q significantly improve the performance of policy learning tasks in noisy environments. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 59f751c0-24ec-4610-99dd-8fada6aef537Cited by top-tier papers3
- Automatic Curriculum Learning With Over-repetition Penalty for Dialogue Policy LearningYangyang Zhao, Zhenyu Wang, Zhenhua HuangAAAI 2021 · 20 citations
- Generative Partial Visual-Tactile Fused Object ClusteringTao Zhang, Yang Cong, Gan Sun, Jiahua Dong et al.AAAI 2021 · 16 citations
- Efficient Dialogue Complementary Policy Learning via Deep Q-network Policy and Episodic Memory PolicyYangyang Zhao, Zhenyu Wang, Changxi Zhu, Shihan WangEMNLP 2021 · 12 citations
Related papers
- Task-Completion Dialogue Policy Learning via Monte Carlo Tree Search with Dueling NetworkSihan Wang, Kaijie Zhou, Kunfeng Lai, Jianping ShenEMNLP 2020 · 10 citations
- Multi-Agent Task-Oriented Dialog Policy Learning with Role-Aware Reward DecompositionRyuichi Takanobu, Runze Liang, Minlie HuangACL 2020 · 47 citations
- Semi-Supervised Dialogue Policy Learning via Stochastic Reward EstimationXinting Huang, Jianzhong Qi, Yu Sun, Rui ZhangACL 2020 · 19 citations
- DyBBT: Dynamic Balance via Bandit-inspired Targeting for Dialog Policy with Cognitive Dual SystemsShuyu Zhang, Yifan Wei, Jialuo Yuan, Xinru Wang et al.ACL 2026
- An Efficient Task-Oriented Dialogue Policy: Evolutionary Reinforcement Learning Injected by Elite IndividualsYangyang Zhao, Ben Niu, Libo Qin, Shihan WangACL 2025 · 3 citations
