A General Highly Accurate Online Planning Method Integrating Large Language Models into Nested Rollout Policy Adaptation for Dialogue Tasks
Hui Wang, Fafa Zhang, Xiaoyu Zhang, Chaoxu Mu
摘要
In goal-oriented dialogue tasks, the main challenge is to steer the interaction towards a given goal within a limited number of turns. Existing approaches either rely on elaborate prompt engineering, whose effectiveness is heavily dependent on human experience, or integrate policy networks and pre-trained policy models, which are usually difficult to adapt to new dialogue scenarios and costly to train. Therefore, in this paper, we present Nested Rollout Policy Adaptation for Goal-oriented Dialogue (NRPA-GD), a novel dialogue policy planning method that completely avoids specific model training by utilizing a Large Language Model (LLM) to simulate behaviors of user and system at the same time. Specifically, NRPA-GD constructs a complete evaluation mechanism for dialogue trajectories and employs an optimization framework of nested Monte Carlo simulation and policy self-adaptation to dynamically adjust policies during the dialogue process. The experimental results on four typical goal-oriented dialogue datasets show that NRPA-GD outperforms both existing prompt engineering and specifically pre-trained model-based methods. Impressively, NRPA-GD surpasses ChatGPT and pre-trained policy models with only a 0.6-billion-parameter LLM. The proposed approach further demonstrates the advantages and novelty of employing planning methods on LLMs to solve practical planning tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree SearchDan Zhang, Sining Zhoubian, Ziniu Hu, Yisong Yue 等NeurIPS 2024 · 被引用 527 次
- Task-Completion Dialogue Policy Learning via Monte Carlo Tree Search with Dueling NetworkSihan Wang, Kaijie Zhou, Kunfeng Lai, Jianping ShenEMNLP 2020 · 被引用 10 次
- Simulating Before Planning: Constructing Intrinsic User World Model for User-Tailored Dialogue Policy PlanningTao He, Lizi Liao, Ming Liu, Bing QinSIGIR 2025 · 被引用 2 次
- Strength Lies in Differences! Improving Strategy Planning for Non-collaborative Dialogues via Diversified User SimulationTong Zhang, Chen Huang, Yang Deng, Hongru Liang 等EMNLP 2024 · 被引用 1 次
- Towards Emotional Support Dialog SystemsSiyang Liu, Chujie Zheng, Orianna Demasi, Sahand Sabour 等ACL 2021
相关 Paper
- Plug-and-Play Policy Planner for Large Language Model Powered Dialogue AgentsYang Deng, Wenxuan Zhang, Wai Lam, See-Kiong Ng 等ICLR 2024 · 被引用 86 次
- Retroformer: Retrospective Large Language Agents with Policy Gradient OptimizationWeiran Yao, Shelby Heinecke, Juan Carlos Niebles, Zhiwei Liu 等ICLR 2024 · 被引用 124 次
- ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue AgentsZhigen Li, Jianxiang Peng, Yanmeng Wang, Yong Cao 等ACL 2025 · 被引用 9 次
- Simulation-Free Hierarchical Latent Policy Planning for Proactive DialoguesTao He, Lizi Liao, Yixin Cao, Yuanxing Liu 等AAAI 2025 · 被引用 11 次
- Planning without Search: Refining Frontier LLMs with Offline Goal-Conditioned RLJoey Hong, Anca D. Dragan, Sergey LevineNeurIPS 2025 · 被引用 11 次
