CoCo: Controllable Counterfactuals for Evaluating Dialogue State Trackers
Shiyang Li, Semih Yavuz, Kazuma Hashimoto, Jia Li, Tong Niu, Nazneen Fatema Rajani, Xifeng Yan, Yingbo Zhou, Caiming Xiong
摘要
Dialogue state trackers have made significant progress on benchmark datasets, but their generalization capability to novel and realistic scenarios beyond the held-out conversations is less understood. We propose controllable counterfactuals (CoCo) to bridge this gap and evaluate dialogue state tracking (DST) models on novel scenarios, i.e., would the system successfully tackle the request if the user responded differently but still consistently with the dialogue flow? CoCo leverages turn-level belief states as counterfactual conditionals to produce novel conversation scenarios in two steps: (i) counterfactual goal generation at turn-level by dropping and adding slots followed by replacing slot values, (ii) counterfactual conversation generation that is conditioned on (i) and consistent with the dialogue flow. Evaluating state-of-the-art DST models on MultiWOZ dataset with CoCo-generated counterfactuals results in a significant performance drop of up to 30.8% (from 49.4% to 18.6%) in absolute joint goal accuracy. In comparison, widely used techniques like paraphrasing only affect the accuracy by at most 2%. Human evaluations show that CoCo-generated conversations perfectly reflect the underlying user goal with more than 95% accuracy and are as human-like as the original conversations, further strengthening its reliability and promise to be adopted as part of the robustness evaluation of DST models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- A Causal Lens for Controllable Text GenerationZhiting Hu, Li Erran LiNeurIPS 2021 · 被引用 77 次
- Unified Dialog Model Pre-training for Task-Oriented Dialog Understanding and GenerationWanwei He, Yinpei Dai, Min Yang, Jian Sun 等SIGIR 2022 · 被引用 41 次
- Large Language Models as Zero-shot Dialogue State Tracker through Function CallingZekun Li, Zhiyu Chen, Mike Ross, Patrick Huber 等ACL 2024 · 被引用 9 次
- Dialogue State Distillation Network with Inter-slot Contrastive Learning for Dialogue State TrackingJing Xu, Dandan Song, Chong Liu, Siu Cheung Hui 等AAAI 2023 · 被引用 8 次
- Divide, Conquer, and Combine: Mixture of Semantic-Independent Experts for Zero-Shot Dialogue State TrackingQingyue Wang, Liang Ding, Yanan Cao, Yibing Zhan 等ACL 2023 · 被引用 5 次
它引用的顶会 Paper9
- Towards Scalable Multi-Domain Conversational Agents: The Schema-Guided Dialogue DatasetAbhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta 等AAAI 2020 · 被引用 707 次
- A Simple Language Model for Task-Oriented DialogueEhsan Hosseini-Asl, Bryan McCann, Chien-Sheng Wu, Semih Yavuz 等NeurIPS 2020 · 被引用 590 次
- Task-Oriented Dialog Systems That Consider Multiple Appropriate Responses under the Same ContextYichi Zhang, Zhijian Ou, Zhou YuAAAI 2020 · 被引用 198 次
- Schema-Guided Multi-Domain Dialogue State Tracking with Graph Attention Neural NetworksLu Chen, Boer Lv, Chi Wang, Su Zhu 等AAAI 2020 · 被引用 143 次
- MinTL: Minimalist Transfer Learning for Task-Oriented Dialogue SystemsZhaojiang Lin, Andrea Madotto, Genta Indra Winata, Pascale FungEMNLP 2020 · 被引用 138 次
相关 Paper
- Correctable-DST: Mitigating Historical Context Mismatch between Training and Inference for Improved Dialogue State TrackingHongyan Xie, Haoxiang Su, Shuangyong Song, Hao Huang 等EMNLP 2022 · 被引用 10 次
- MetaASSIST: Robust Dialogue State Tracking with Meta LearningFanghua Ye, Xi Wang, Jie Huang, Shenghui Li 等EMNLP 2022 · 被引用 10 次
- BREAK: Breaking the Dialogue State Tracking Barrier with Beam Search and Re-rankingSeungpil Won, Heeyoung Kwak, Joongbo Shin, Janghoon Han 等ACL 2023 · 被引用 5 次
- Multi-domain Dialogue State Tracking with Recursive InferenceLizi Liao, Tongyao Zhu, Le Hong Long, Tat-Seng ChuaWWW 2021 · 被引用 11 次
- Dual Slot Selector via Local Reliability Verification for Dialogue State TrackingJinyu Guo, Kai Shuang, Jijie Li, Zihan WangACL 2021
