Offline Reinforcement Learning for Mixture-of-Expert Dialogue Management
Dhawal Gupta, Yinlam Chow, Azamat Tulepbergenov, Mohammad Ghavamzadeh, Craig Boutilier
摘要
Reinforcement learning (RL) has shown great promise for developing agents for dialogue management (DM) that are non-myopic, conduct rich conversations, and maximize overall user satisfaction. Despite the advancements in RL and language models (LMs), employing RL to drive conversational chatbots still poses significant challenges. A primary issue stems from RL's dependency on online exploration for effective learning, a process that can be costly. Moreover, engaging in online interactions with humans during the training phase can raise safety concerns, as the LM can potentially generate unwanted outputs. This issue is exacerbated by the combinatorial action spaces facing these algorithms, as most LM agents generate responses at the word level. We develop various RL algorithms, specialized in dialogue planning, that leverage recent Mixture-of-Expert Language Models (MoE-LMs)-models that capture diverse semantics, generate utterances reflecting different intents, and are amenable for multi-turn DM. By exploiting the MoE-LM structure, our methods significantly reduce the size of the action space and improve the efficacy of RL-based DM. We evaluate our methods in open-domain dialogue to demonstrate their effectiveness with respect to the diversity of intent in generated utterances and overall DM performance. * The work was done as a student researcher at Google Research
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- Hierarchical Reinforcement Learning for Open-Domain DialogAbdelrhman Saleh, Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen 等AAAI 2020 · 被引用 60 次
- Offline RL for Natural Language Generation with Implicit Language Q LearningCharlie Snell, Ilya Kostrikov, Yi Su, Sherry Yang 等ICLR 2023 · 被引用 9 次
- A Mixture-of-Expert Approach to RL-based Dialogue ManagementYinlam Chow, Aza Tulepbergenov, Ofir Nachum, Dhawal Gupta 等ICLR 2023 · 被引用 2 次
相关 Paper
- Human-centric dialog training via offline reinforcement learningNatasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun, Craig Ferguson 等EMNLP 2020 · 被引用 9 次
- Mixture-of-Minds: Multi-Agent Reinforcement Learning for Table UnderstandingYuhang Zhou, Mingrui Zhang, Ke Li, Mingyi Wang 等ACL 2026 · 被引用 5 次
- Knowledge Graph Grounded Goal Planning for Open-Domain Conversation GenerationJun Xu, Haifeng Wang, Zhengyu Niu, Hua Wu 等AAAI 2020 · 被引用 69 次
- [CASPI] Causal-aware Safe Policy Improvement for Task-oriented DialogueGovardana Sachithanandam Ramachandran, Kazuma Hashimoto, Caiming XiongACL 2022 · 被引用 12 次
- Planning Like Human: A Dual-process Framework for Dialogue PlanningTao He, Lizi Liao, Yixin Cao, Yuanxing Liu 等ACL 2024 · 被引用 7 次
