Bayes-Adaptive Monte-Carlo Planning and Learning for Goal-Oriented Dialogues
Youngsoo Jang, Jongmin Lee, Kee-Eung Kim
Abstract
We consider a strategic dialogue task, where the ability to infer the other agent's goal is critical to the success of the conversational agent. While this problem can be naturally formulated as Bayesian planning, it is known to be a very difficult problem due to its enormous search space consisting of all possible utterances. In this paper, we introduce an efficient Bayes-adaptive planning algorithm for goal-oriented dialogues, which combines RNN-based dialogue generation and MCTS-based Bayesian planning in a novel way, leading to robust decision-making under the uncertainty of the other agent's goal. We then introduce reinforcement learning for the dialogue agent that uses MCTS as a strong policy improvement operator, casting reinforcement learning as iterative alternation of planning and supervised-learning of self-generated dialogues. In the experiments, we demonstrate that our Bayes-adaptive dialogue planning agent significantly outperforms the state-of-the-art in a negotiation dialogue domain. We also show that reinforcement learning via MCTS further improves end-task performance without diverging from human language. under the uncertainty of the other agent's goal. While this can be naturally formulated as Bayesian planning, computing Bayes-optimal policy itself is generally infeasible except for very small-scale problems. Second, optimizing the agent through goal-based training by vanilla reinforcement learning (e.g. REINFORCE) is inefficient and unstable due to the high variance of policy gradient estimator, and it typically leads to divergence from human language (Lewis et al. 2017; Buck et al. 2018) . Due to the inherent difficulty of Bayesian planning, existing works for the end-to-end goal-based dialogue agent either do not perform multi-step planning or just adopt a simple dialogue rollout with an arbitrarily fixed goal of the
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4bbb679b-1bdb-4bbc-9dee-0c67061c2447Cited by top-tier papers5
- Planning Like Human: A Dual-process Framework for Dialogue PlanningTao He, Lizi Liao, Yixin Cao, Yuanxing Liu et al.ACL 2024 · 7 citations
- Strength Lies in Differences! Improving Strategy Planning for Non-collaborative Dialogues via Diversified User SimulationTong Zhang, Chen Huang, Yang Deng, Hongru Liang et al.EMNLP 2024 · 1 citation
- Degeneration-free Policy Optimization: RL Fine-Tuning for Language Models without DegenerationYoungsoo Jang, Geon-Hyeong Kim, Byoungjip Kim, Yu Jin Kim et al.ICML 2024 · 1 citation
- Improving Dialog Systems for Negotiation with Personality ModelingRunzhe Yang, Jingxiao Chen, Karthik NarasimhanACL 2021
- Symbolic Planning and Code Generation for Grounded DialogueJustin T. Chiu, Wenting Zhao, Derek Chen, Saujas Vaduguru et al.EMNLP 2023
Related papers
- Task-Completion Dialogue Policy Learning via Monte Carlo Tree Search with Dueling NetworkSihan Wang, Kaijie Zhou, Kunfeng Lai, Jianping ShenEMNLP 2020 · 10 citations
- Online Bayesian Goal Inference for Boundedly Rational Planning AgentsTan Zhi-Xuan, Jordyn L. Mann, Tom Silver, Josh Tenenbaum et al.NeurIPS 2020 · 122 citations
- EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement LearningXiaoqian Liu, Ke Wang, Yongbin Li, Yuchuan Wu et al.ACL 2025 · 7 citations
- Meta-Reinforced Multi-Domain State Generator for Dialogue SystemsYi Huang, Junlan Feng, Min Hu, Xiaoting Wu et al.ACL 2020 · 30 citations
- A Mixture-of-Expert Approach to RL-based Dialogue ManagementYinlam Chow, Aza Tulepbergenov, Ofir Nachum, Dhawal Gupta et al.ICLR 2023 · 2 citations
