Uncertainty Measures in Neural Belief Tracking and the Effects on Dialogue Policy Performance
Carel van Niekerk, Andrey Malinin, Christian Geishauser, Michael Heck, Hsien-Chin Lin, Nurul Lubis, Shutong Feng, Milica Gasic
Abstract
The ability to identify and resolve uncertainty is crucial for the robustness of a dialogue system. Indeed, this has been confirmed empirically on systems that utilise Bayesian approaches to dialogue belief tracking. However, such systems consider only confidence estimates and have difficulty scaling to more complex settings. Neural dialogue systems, on the other hand, rarely take uncertainties into account. They are therefore overconfident in their decisions and less robust. Moreover, the performance of the tracking task is often evaluated in isolation, without consideration of its effect on the downstream policy optimisation. We propose the use of different uncertainty measures in neural belief tracking. The effects of these measures on the downstream task of policy optimisation are evaluated by adding selected measures of uncertainty to the feature space of the policy and training policies through interaction with a user simulator. Both human and simulated user results show that incorporating these measures leads to improvements both of the performance and of the robustness of the downstream dialogue policy. This highlights the importance of developing neural dialogue belief trackers that take uncertainty into account.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9a72177a-5a80-421e-b375-1bd4c0ffb988Builds on5
- Pitfalls of In-Domain Uncertainty Estimation and Ensembling in Deep LearningArsenii Ashukha, Alexander Lyzhov, Dmitry Molchanov, Dmitry P. VetrovICLR 2020 · 354 citations
- Ensemble Distribution DistillationAndrey Malinin, Bruno Mlodozeniec, Mark J. F. GalesICLR 2020 · 273 citations
- Slot Self-Attentive Dialogue State TrackingFanghua Ye, Jarana Manotumruksa, Qiang Zhang, Shenghui Li et al.WWW 2021 · 68 citations
- SAS: Dialogue State Tracking via Slot Attention and Slot Information SharingJiaying Hu, Yan Yang, Chencai Chen, Liang He et al.ACL 2020 · 46 citations
- Scaling Ensemble Distribution Distillation to Many Classes with Proxy TargetsMax Ryabinin, Andrey Malinin, Mark J. F. GalesNeurIPS 2021 · 23 citations
Related papers
- Evaluating Theory of (an uncertain) Mind: Predicting the Uncertain Beliefs of Others from Conversational CuesAnthony B. Sicilia, Malihe AlikhaniACL 2025 · 2 citations
- On the Calibration and Uncertainty with Pólya-Gamma Augmentation for Dialog Retrieval ModelsTong Ye, Shijing Si, Jianzong Wang, Ning Cheng et al.AAAI 2023 · 3 citations
- The Neural Testbed: Evaluating Joint PredictionsIan Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla et al.NeurIPS 2022 · 28 citations
- Beyond Expected Return: Accounting for Policy Reproducibility When Evaluating Reinforcement Learning AlgorithmsManon Flageat, Bryan Lim, Antoine CullyAAAI 2024 · 4 citations
- CoCo: Controllable Counterfactuals for Evaluating Dialogue State TrackersShiyang Li, Semih Yavuz, Kazuma Hashimoto, Jia Li et al.ICLR 2021 · 65 citations
