Don't be Contradicted with Anything! CI-ToD: Towards Benchmarking Consistency for Task-oriented Dialogue System
Libo Qin, Tianbao Xie, Shijue Huang, Qiguang Chen, Xiao Xu, Wanxiang Che
摘要
Consistency Identification has obtained remarkable success on open-domain dialogue, which can be used for preventing inconsistent response generation. However, in contrast to the rapid development in open-domain dialogue, few efforts have been made to the task-oriented dialogue direction. In this paper, we argue that consistency problem is more urgent in task-oriented domain. To facilitate the research, we introduce CI-ToD, a novel dataset for Consistency Identification in Taskoriented Dialog system. In addition, we not only annotate the single label to enable the model to judge whether the system response is contradictory, but also provide more finegrained labels (i.e., Dialogue History Inconsistency, User Query Inconsistency and Knowledge Base Inconsistency) to encourage model to know what inconsistent sources lead to it. Empirical results show that state-of-the-art methods only achieve 51.3%, which is far behind the human performance of 93.2%, indicating that there is ample room for improving consistency identification ability. Finally, we conduct exhaustive experiments and qualitative analysis to comprehend key challenges and provide guidance for future directions. All datasets and models are publicly available at https://github.com/yizhen20133868/CI-ToD . * Email corresponding. User: Give me directions to the closest grocery store. System: There is a whole foods 2 miles away and their address is 880_ames_ct. User: I need a route that avoids all heavy traffic.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- CDConv: A Benchmark for Contradiction Detection in Chinese ConversationsChujie Zheng, Jinfeng Zhou, Yinhe Zheng, Libiao Peng 等EMNLP 2022 · 被引用 6 次
- Instruct Once, Chat Consistently in Multiple Rounds: An Efficient Tuning Framework for DialogueJian Wang, Chak Tou Leong, Jiashuo Wang, Dongding Lin 等ACL 2024 · 被引用 4 次
- Red Teaming Language Models for Processing Contradictory DialoguesXiaofei Wen, Bangzheng Li, Tenghao Huang, Muhao ChenEMNLP 2024 · 被引用 1 次
- When Misinformation Speaks and Converses: Rethinking Fact-Checking in Audio PlatformsChaewan Chun, Delvin Ce Zhang, Dongwon LeeACL 2026
- DialFact: A Benchmark for Fact-Checking in DialoguePrakhar Gupta, Chien-Sheng Wu, Wenhao Liu, Caiming XiongACL 2022
它引用的顶会 Paper9
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang 等ICLR 2020 · 被引用 674 次
- DCR-Net: A Deep Co-Interactive Relation Network for Joint Dialog Act Recognition and Sentiment ClassificationLibo Qin, Wanxiang Che, Yangming Li, Minheng Ni 等AAAI 2020 · 被引用 100 次
- Dynamic Fusion Network for Multi-Domain End-to-end Task-Oriented DialogLibo Qin, Xiao Xu, Wanxiang Che, Yue Zhang 等ACL 2020 · 被引用 90 次
- Multi-Agent Task-Oriented Dialog Policy Learning with Role-Aware Reward DecompositionRyuichi Takanobu, Runze Liang, Minlie HuangACL 2020 · 被引用 47 次
相关 Paper
- Q-TOD: A Query-driven Task-oriented Dialogue SystemXin Tian, Yingzhan Lin, Mengfei Song, Siqi Bao 等EMNLP 2022 · 被引用 13 次
- I like fish, especially dolphins: Addressing Contradictions in Dialogue ModelingYixin Nie, Mary Williamson, Mohit Bansal, Douwe Kiela 等ACL 2021
- Enhancing Goal-oriented Proactive Dialogue Systems via Consistency Reflection and CorrectionDidi Zhang, Yaxin Fan, Peifeng Li, Qiaoming ZhuACL 2025 · 被引用 1 次
- From Retrieval to Generation: A Simple and Unified Generative Model for End-to-End Task-Oriented DialogueZeyuan Ding, Zhihao Yang, Ling Luo, Yuanyuan Sun 等AAAI 2024 · 被引用 6 次
- TOD-BERT: Pre-trained Natural Language Understanding for Task-Oriented DialogueChien-Sheng Wu, Steven C. H. Hoi, Richard Socher, Caiming XiongEMNLP 2020 · 被引用 210 次
