Uncertainty Measures in Neural Belief Tracking and the Effects on Dialogue Policy Performance
Carel van Niekerk, Andrey Malinin, Christian Geishauser, Michael Heck, Hsien-Chin Lin, Nurul Lubis, Shutong Feng, Milica Gasic
摘要
The ability to identify and resolve uncertainty is crucial for the robustness of a dialogue system. Indeed, this has been confirmed empirically on systems that utilise Bayesian approaches to dialogue belief tracking. However, such systems consider only confidence estimates and have difficulty scaling to more complex settings. Neural dialogue systems, on the other hand, rarely take uncertainties into account. They are therefore overconfident in their decisions and less robust. Moreover, the performance of the tracking task is often evaluated in isolation, without consideration of its effect on the downstream policy optimisation. We propose the use of different uncertainty measures in neural belief tracking. The effects of these measures on the downstream task of policy optimisation are evaluated by adding selected measures of uncertainty to the feature space of the policy and training policies through interaction with a user simulator. Both human and simulated user results show that incorporating these measures leads to improvements both of the performance and of the robustness of the downstream dialogue policy. This highlights the importance of developing neural dialogue belief trackers that take uncertainty into account.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Pitfalls of In-Domain Uncertainty Estimation and Ensembling in Deep LearningArsenii Ashukha, Alexander Lyzhov, Dmitry Molchanov, Dmitry P. VetrovICLR 2020 · 被引用 354 次
- Ensemble Distribution DistillationAndrey Malinin, Bruno Mlodozeniec, Mark J. F. GalesICLR 2020 · 被引用 273 次
- Slot Self-Attentive Dialogue State TrackingFanghua Ye, Jarana Manotumruksa, Qiang Zhang, Shenghui Li 等WWW 2021 · 被引用 68 次
- SAS: Dialogue State Tracking via Slot Attention and Slot Information SharingJiaying Hu, Yan Yang, Chencai Chen, Liang He 等ACL 2020 · 被引用 46 次
- Scaling Ensemble Distribution Distillation to Many Classes with Proxy TargetsMax Ryabinin, Andrey Malinin, Mark J. F. GalesNeurIPS 2021 · 被引用 23 次
相关 Paper
- Evaluating Theory of (an uncertain) Mind: Predicting the Uncertain Beliefs of Others from Conversational CuesAnthony B. Sicilia, Malihe AlikhaniACL 2025 · 被引用 2 次
- On the Calibration and Uncertainty with Pólya-Gamma Augmentation for Dialog Retrieval ModelsTong Ye, Shijing Si, Jianzong Wang, Ning Cheng 等AAAI 2023 · 被引用 3 次
- The Neural Testbed: Evaluating Joint PredictionsIan Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla 等NeurIPS 2022 · 被引用 28 次
- Beyond Expected Return: Accounting for Policy Reproducibility When Evaluating Reinforcement Learning AlgorithmsManon Flageat, Bryan Lim, Antoine CullyAAAI 2024 · 被引用 4 次
- CoCo: Controllable Counterfactuals for Evaluating Dialogue State TrackersShiyang Li, Semih Yavuz, Kazuma Hashimoto, Jia Li 等ICLR 2021 · 被引用 65 次
