Planning Like Human: A Dual-process Framework for Dialogue Planning
Tao He, Lizi Liao, Yixin Cao, Yuanxing Liu, Ming Liu, Zerui Chen, Bing Qin
摘要
In proactive dialogue, the challenge lies not just in generating responses but in steering conversations toward predetermined goals, a task where Large Language Models (LLMs) typically struggle due to their reactive nature. Traditional approaches to enhance dialogue planning in LLMs, ranging from elaborate prompt engineering to the integration of policy networks, either face efficiency issues or deliver suboptimal performance. Inspired by the dualprocess theory in psychology, which identifies two distinct modes of thinking-intuitive (fast) and analytical (slow), we propose the Dual-Process Dialogue Planning (DPDP) framework. DPDP embodies this theory through two complementary planning systems: an instinctive policy model for familiar contexts and a deliberative Monte Carlo Tree Search (MCTS) mechanism for complex, novel scenarios. This dual strategy is further coupled with a novel two-stage training regimen: offline Reinforcement Learning for robust initial policy model formation followed by MCTS-enhanced on-thefly learning, which ensures a dynamic balance between efficiency and strategic depth. Our empirical evaluations across diverse dialogue tasks affirm DPDP's superiority in achieving both high-quality dialogues and operational efficiency, outpacing existing methods. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Adaptive Social Learning via Mode Policy Optimization for Language AgentsMinzheng Wang, Yongbin Li, Haobo Wang, Xinghua Zhang 等ICLR 2026 · 被引用 15 次
- Leveraging Dual Process Theory in Language Agent Framework for Real-time Simultaneous Human-AI CollaborationShao Zhang, Xihuai Wang, Wenhao Zhang, Chaoran Li 等ACL 2025 · 被引用 14 次
- Simulation-Free Hierarchical Latent Policy Planning for Proactive DialoguesTao He, Lizi Liao, Yixin Cao, Yuanxing Liu 等AAAI 2025 · 被引用 11 次
- EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement LearningXiaoqian Liu, Ke Wang, Yongbin Li, Yuchuan Wu 等ACL 2025 · 被引用 7 次
- A Dual-Mind Framework for Strategic and Expressive Negotiation AgentYutong Liu, Lida Shi, Rui Song, Hao XuACL 2025 · 被引用 4 次
它引用的顶会 Paper13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Guiding Large Language Models via Directional Stimulus PromptingZekun Li, Baolin Peng, Pengcheng He, Michel Galley 等NeurIPS 2023 · 被引用 163 次
- Learning to summarize with human feedbackNisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler 等NeurIPS 2020 · 被引用 124 次
- Retroformer: Retrospective Large Language Agents with Policy Gradient OptimizationWeiran Yao, Shelby Heinecke, Juan Carlos Niebles, Zhiwei Liu 等ICLR 2024 · 被引用 124 次
相关 Paper
- Plug-and-Play Policy Planner for Large Language Model Powered Dialogue AgentsYang Deng, Wenxuan Zhang, Wai Lam, See-Kiong Ng 等ICLR 2024 · 被引用 86 次
- Thinker: Learning to Think Fast and SlowStephen Chung, Wenyu Du, Jie FuNeurIPS 2025 · 被引用 10 次
- SwiftSage: A Generative Agent with Fast and Slow Thinking for Complex Interactive TasksBill Yuchen Lin, Yicheng Fu, Karina Yang, Faeze Brahman 等NeurIPS 2023 · 被引用 244 次
- ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue AgentsZhigen Li, Jianxiang Peng, Yanmeng Wang, Yong Cao 等ACL 2025 · 被引用 9 次
- A General Highly Accurate Online Planning Method Integrating Large Language Models into Nested Rollout Policy Adaptation for Dialogue TasksHui Wang, Fafa Zhang, Xiaoyu Zhang, Chaoxu MuAAAI 2026
