An Efficient Task-Oriented Dialogue Policy: Evolutionary Reinforcement Learning Injected by Elite Individuals
Yangyang Zhao, Ben Niu, Libo Qin, Shihan Wang
摘要
Deep Reinforcement Learning (DRL) is widely used in task-oriented dialogue systems to optimize dialogue policy, but it struggles to balance exploration and exploitation due to the high dimensionality of state and action spaces. This challenge often results in local optima or poor convergence. Evolutionary Algorithms (EAs) have been proven to effectively explore the solution space of neural networks by maintaining population diversity. Inspired by this, we innovatively combine the global search capabilities of EA with the local optimization of DRL to achieve a balance between exploration and exploitation. Nevertheless, the inherent flexibility of natural language in dialogue tasks complicates this direct integration, leading to prolonged evolutionary times. Thus, we further propose an elite individual injection mechanism to enhance EA's search efficiency by adaptively introducing best-performing individuals into the population. Experiments across four datasets show that our approach significantly improves the balance between exploration and exploitation, boosting performance. Moreover, the effectiveness of the EII mechanism in reducing exploration time has been demonstrated, achieving an efficient integration of EA and DRL on task-oriented dialogue policy tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Benchmarking and Learning Real-World Customer Service DialogueTianhong Gao, Jundong Shen, Jiapeng Wang, Bei Shi 等ACL 2026
- DyBBT: Dynamic Balance via Bandit-inspired Targeting for Dialog Policy with Cognitive Dual SystemsShuyu Zhang, Yifan Wei, Jialuo Yuan, Xinru Wang 等ACL 2026
它引用的顶会 Paper5
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Learning Efficient Dialogue Policy from Demonstrations through ShapingHuimin Wang, Baolin Peng, Kam-Fai WongACL 2020 · 被引用 18 次
- ERL-Re: Efficient Evolutionary Reinforcement Learning with Shared State Representation and Individual Policy RepresentationJianye Hao, Pengyi Li, Hongyao Tang, Yan Zheng 等ICLR 2023 · 被引用 16 次
- End-to-end Task-oriented Dialogue: A Survey of Tasks, Methods, and Future DirectionsLibo Qin, Wenbo Pan, Qiguang Chen, Lizi Liao 等EMNLP 2023 · 被引用 12 次
- Value-Evolutionary-Based Reinforcement LearningPengyi Li, Jianye Hao, Hongyao Tang, Yan Zheng 等ICML 2024 · 被引用 10 次
相关 Paper
- EvoRainbow: Combining Improvements in Evolutionary Reinforcement Learning for Policy SearchPengyi Li, Yan Zheng, Hongyao Tang, Xian Fu 等ICML 2024 · 被引用 13 次
- Cooperative Heterogeneous Deep Reinforcement LearningHan Zheng, Pengfei Wei, Jing Jiang, Guodong Long 等NeurIPS 2020 · 被引用 20 次
- An Efficient Asynchronous Method for Integrating Evolutionary and Gradient-based Policy SearchKyunghyun Lee, Byeong-Uk Lee, Ukcheol Shin, In So KweonNeurIPS 2020 · 被引用 24 次
- Two-Stage Evolutionary Reinforcement Learning for Enhancing Exploration and ExploitationQingling Zhu, Xiaoqiang Wu, Qiuzhen Lin, Wei-Neng ChenAAAI 2024 · 被引用 9 次
- DarwinTOD: LLM-Driven Lifelong Self-evolution for Task-oriented Dialog SystemsShuyu Zhang, Yujie Liu, Xinru Wang, Cheng Zhang 等ACL 2026 · 被引用 1 次
