DyBBT: Dynamic Balance via Bandit-inspired Targeting for Dialog Policy with Cognitive Dual Systems
Shuyu Zhang, Yifan Wei, Jialuo Yuan, Xinru Wang, Yanmin Zhu, Yujie Liu, Bin Li
Abstract
Task oriented dialog systems often rely on static exploration strategies that do not adapt to dynamic dialog contexts, leading to inefficient exploration and suboptimal performance. We propose DyBBT 1 , a novel dialog policy learning framework that formalizes the exploration challenge through a structured cognitive state space C that captures dialog progression, user uncertainty, and slot dependency. DyBBT proposes a bandit-inspired meta-controller that dynamically switches between a fast intuitive inference (System 1) and a slow deliberative reasoner (System 2) based on real-time cognitive states and visitation counts. Extensive experiments on single-and multi-domain benchmarks show that DyBBT achieves SOTA performance in success rate, efficiency, and generalization, with human evaluations confirming that its decisions are well-aligned with expert judgment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ed2d1668-450e-42a3-90a0-e38c907cb8cfBuilds on13
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Beyond UCB: Optimal and Efficient Contextual Bandits with Regression OraclesDylan J. Foster, Alexander RakhlinICML 2020 · 241 citations
- GALAXY: A Generative Pre-trained Model for Task-Oriented Dialog with Semi-supervised Learning and Explicit Policy InjectionWanwei He, Yinpei Dai, Yinhe Zheng, Yuchuan Wu et al.AAAI 2022 · 181 citations
- Towards Human-AI Deliberation: Design and Evaluation of LLM-Empowered Deliberative AI for AI-Assisted Decision-MakingShuai Ma, Qiaoyi Chen, Xinru Wang, Chengbo Zheng et al.CHI 2025 · 113 citations
- Bridging the Gap between Prior and Posterior Knowledge Selection for Knowledge-Grounded Dialogue GenerationXiuyi Chen, Fandong Meng, Peng Li, Feilong Chen et al.EMNLP 2020 · 78 citations
Related papers
- Dynamic Reward-Based Dueling Deep Dyna-Q: Robust Policy Learning in Noisy EnvironmentsYangyang Zhao, Zhenyu Wang, Kai Yin, Rui Zhang et al.AAAI 2020 · 32 citations
- Task-Completion Dialogue Policy Learning via Monte Carlo Tree Search with Dueling NetworkSihan Wang, Kaijie Zhou, Kunfeng Lai, Jianping ShenEMNLP 2020 · 10 citations
- Meta-Reinforced Multi-Domain State Generator for Dialogue SystemsYi Huang, Junlan Feng, Min Hu, Xiaoting Wu et al.ACL 2020 · 30 citations
- Exploring Auxiliary Reasoning Tasks for Task-oriented Dialog Systems with Meta Cooperative LearningBowen Qin, Min Yang, Lidong Bing, Qingshan Jiang et al.AAAI 2021 · 9 citations
- ChatR1: Reinforcement Learning for Conversational Reasoning and Retrieval Augmented Question AnsweringSimon Lupart, Mohammad Aliannejadi, Evangelos KanoulasACL 2026 · 5 citations
