Automatic Curriculum Learning With Over-repetition Penalty for Dialogue Policy Learning
Yangyang Zhao, Zhenyu Wang, Zhenhua Huang
摘要
Dialogue policy learning based on reinforcement learning is difficult to be applied to real users to train dialogue agents from scratch because of the high cost. User simulators, which choose random user goals for the dialogue agent to train on, have been considered as an affordable substitute for real users. However, this random sampling method ignores the law of human learning, making the learned dialogue policy inefficient and unstable. We propose a novel framework, Automatic Curriculum Learning-based Deep Q-Network (ACL-DQN), which replaces the traditional random sampling method with a teacher policy model to realize the dialogue policy for automatic curriculum learning. The teacher model arranges a meaningful ordered curriculum and automatically adjusts it by monitoring the learning progress of the dialogue agent and the over-repetition penalty without any requirement of prior knowledge. The learning progress of the dialogue agent reflects the relationship between the dialogue agent's ability and the sampled goals' difficulty for sample efficiency. The over-repetition penalty guarantees the sampled diversity. Experiments show that the ACL-DQN significantly improves the effectiveness and stability of dialogue tasks with a statistically significant margin. Furthermore, the framework can be further improved by equipping with different curriculum schedules, which demonstrates that the framework has strong generalizability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Generative Partial Visual-Tactile Fused Object ClusteringTao Zhang, Yang Cong, Gan Sun, Jiahua Dong 等AAAI 2021 · 被引用 16 次
- Efficient Dialogue Complementary Policy Learning via Deep Q-network Policy and Episodic Memory PolicyYangyang Zhao, Zhenyu Wang, Changxi Zhu, Shihan WangEMNLP 2021 · 被引用 12 次
- PMG: Progressive Motion Generation via Sparse Anchor Postures Curriculum LearningYingjie Xi, Jian Jun Zhang, Xiaosong YangACM MM 2025 · 被引用 1 次
- Bootstrapped Policy Learning for Task-oriented Dialogue through Goal ShapingYangyang Zhao, Ben Niu, Mehdi Dastani, Shihan WangEMNLP 2024
- LaMDAgent: An Autonomous Framework for Post-Training Pipeline Optimization via LLM AgentsTaro Yano, Yoichi Ishibashi, Masafumi OyamadaEMNLP 2025
它引用的顶会 Paper2
相关 Paper
- Efficient Dialog Policy Learning by Reasoning with Contextual KnowledgeHaodi Zhang, Zhichao Zeng, Keting Lu, Kaishun Wu 等AAAI 2022 · 被引用 14 次
- Multi-Agent Task-Oriented Dialog Policy Learning with Role-Aware Reward DecompositionRyuichi Takanobu, Runze Liang, Minlie HuangACL 2020 · 被引用 47 次
- Learning from Easy to Complex: Adaptive Multi-Curricula Learning for Neural Dialogue GenerationHengyi Cai, Hongshen Chen, Cheng Zhang, Yonghao Song 等AAAI 2020 · 被引用 22 次
- Self-Paced Deep Reinforcement LearningPascal Klink, Carlo D'Eramo, Jan Peters, Joni PajarinenNeurIPS 2020 · 被引用 83 次
- Task-Completion Dialogue Policy Learning via Monte Carlo Tree Search with Dueling NetworkSihan Wang, Kaijie Zhou, Kunfeng Lai, Jianping ShenEMNLP 2020 · 被引用 10 次
