One Planner To Guide Them All ! Learning Adaptive Conversational Planners for Goal-oriented Dialogues
Huy Quang Dao, Lizi Liao
摘要
Goal-oriented dialogues, such as recommendation and negotiation, often require balancing multiple conflicting objectives.Conventional approaches typically train separate policies for each predefined objective trade-off, which is computationally costly and scales poorly.In this work, we pursue a single dialogue policy that can dynamically adapt to varying objective preferences at inference time without retraining.This raises several challenges in terms of both (1) optimization strategy and (2) knowledge utilization.To address these, we propose a novel policy learning framework, Preference Adaptive Dialogue Policy Planner (PADPP), for multi-objective goal-oriented dialogues.Specifically, to tackle the former, we introduce a novel optimization scheme, which leverages information gained from training the model on previously updated objective weights, accelerating the learning capability on new weight settings.To address the latter, we utilize Generalized Policy Improvement (GPI) to ensure the effectiveness of leveraged knowledge.Experimental results demonstrate that PADPP achieves superior adaptability and performance compared to state-of-the-art approaches, offering a scalable and flexible solution for multiobjective, goal-oriented dialogues 1 . InferenceRoBERTa DQN LM Planner Action 1 (, , ) Action 2 (, , ) Recommendation Dialogue: DuRecDial 2.0
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Towards Conversational Recommendation over Multi-Type DialogsZeming Liu, Haifeng Wang, Zheng-Yu Niu, Hua Wu 等ACL 2020 · 被引用 157 次
- DuRecDial 2.0: A Bilingual Parallel Corpus for Conversational RecommendationZeming Liu, Haifeng Wang, Zhengyu Niu, Hua Wu 等EMNLP 2021 · 被引用 39 次
- Efficient Dialogue Complementary Policy Learning via Deep Q-network Policy and Episodic Memory PolicyYangyang Zhao, Zhenyu Wang, Changxi Zhu, Shihan WangEMNLP 2021 · 被引用 12 次
- Task-Completion Dialogue Policy Learning via Monte Carlo Tree Search with Dueling NetworkSihan Wang, Kaijie Zhou, Kunfeng Lai, Jianping ShenEMNLP 2020 · 被引用 10 次
- Interacting with Non-Cooperative User: A New Paradigm for Proactive Dialogue PolicyWenqiang Lei, Yao Zhang, Feifan Song, Hongru Liang 等SIGIR 2022 · 被引用 7 次
相关 Paper
- AMOR: Adaptive Character Control through Multi-Objective Reinforcement LearningLucas N. Alegre, Agon Serifi, Ruben Grandia, David Müller 等SIGGRAPH 2025 · 被引用 4 次
- Pareto Continual Learning: Preference-Conditioned Learning and Adaption for Dynamic Stability-Plasticity Trade-offSong Lai, Zhe Zhao, Fei Zhu, Xi Lin 等AAAI 2025 · 被引用 4 次
- Plug-and-Play Policy Planner for Large Language Model Powered Dialogue AgentsYang Deng, Wenxuan Zhang, Wai Lam, See-Kiong Ng 等ICLR 2024 · 被引用 86 次
- A distributional view on multi-objective policy optimizationAbbas Abdolmaleki, Sandy H. Huang, Leonard Hasenclever, Michael Neunert 等ICML 2020 · 被引用 93 次
- Simulation-Free Hierarchical Latent Policy Planning for Proactive DialoguesTao He, Lizi Liao, Yixin Cao, Yuanxing Liu 等AAAI 2025 · 被引用 11 次
