Offline Reinforcement Learning for Mixture-of-Expert Dialogue Management
Dhawal Gupta, Yinlam Chow, Azamat Tulepbergenov, Mohammad Ghavamzadeh, Craig Boutilier
Abstract
Reinforcement learning (RL) has shown great promise for developing agents for dialogue management (DM) that are non-myopic, conduct rich conversations, and maximize overall user satisfaction. Despite the advancements in RL and language models (LMs), employing RL to drive conversational chatbots still poses significant challenges. A primary issue stems from RL's dependency on online exploration for effective learning, a process that can be costly. Moreover, engaging in online interactions with humans during the training phase can raise safety concerns, as the LM can potentially generate unwanted outputs. This issue is exacerbated by the combinatorial action spaces facing these algorithms, as most LM agents generate responses at the word level. We develop various RL algorithms, specialized in dialogue planning, that leverage recent Mixture-of-Expert Language Models (MoE-LMs)-models that capture diverse semantics, generate utterances reflecting different intents, and are amenable for multi-turn DM. By exploiting the MoE-LM structure, our methods significantly reduce the size of the action space and improve the efficacy of RL-based DM. We evaluate our methods in open-domain dialogue to demonstrate their effectiveness with respect to the diversity of intent in generated utterances and overall DM performance. * The work was done as a student researcher at Google Research
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5c26fbce-dea3-4c99-b494-dfcea9fb3529Builds on4
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- Hierarchical Reinforcement Learning for Open-Domain DialogAbdelrhman Saleh, Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen et al.AAAI 2020 · 60 citations
- Offline RL for Natural Language Generation with Implicit Language Q LearningCharlie Snell, Ilya Kostrikov, Yi Su, Sherry Yang et al.ICLR 2023 · 9 citations
- A Mixture-of-Expert Approach to RL-based Dialogue ManagementYinlam Chow, Aza Tulepbergenov, Ofir Nachum, Dhawal Gupta et al.ICLR 2023 · 2 citations
Related papers
- Human-centric dialog training via offline reinforcement learningNatasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun, Craig Ferguson et al.EMNLP 2020 · 9 citations
- Mixture-of-Minds: Multi-Agent Reinforcement Learning for Table UnderstandingYuhang Zhou, Mingrui Zhang, Ke Li, Mingyi Wang et al.ACL 2026 · 5 citations
- Knowledge Graph Grounded Goal Planning for Open-Domain Conversation GenerationJun Xu, Haifeng Wang, Zhengyu Niu, Hua Wu et al.AAAI 2020 · 69 citations
- [CASPI] Causal-aware Safe Policy Improvement for Task-oriented DialogueGovardana Sachithanandam Ramachandran, Kazuma Hashimoto, Caiming XiongACL 2022 · 12 citations
- Planning Like Human: A Dual-process Framework for Dialogue PlanningTao He, Lizi Liao, Yixin Cao, Yuanxing Liu et al.ACL 2024 · 7 citations
