DialFact: A Benchmark for Fact-Checking in Dialogue
Prakhar Gupta, Chien-Sheng Wu, Wenhao Liu, Caiming Xiong
摘要
Fact-checking is an essential tool to mitigate the spread of misinformation and disinformation. We introduce the task of fact-checking in dialogue, which is a relatively unexplored area. We construct DIALFACT, a testing benchmark dataset of 22,245 annotated conversational claims, paired with pieces of evidence from Wikipedia. There are three sub-tasks in DIALFACT: 1) Verifiable claim detection task distinguishes whether a response carries verifiable factual information; 2) Evidence retrieval task retrieves the most relevant Wikipedia snippets as evidence; 3) Claim verification task predicts a dialogue response to be supported, refuted, or not enough information. We found that existing fact-checking models trained on non-dialogue data like FEVER (Thorne et al., 2018) fail to perform well on our task, and thus, we propose a simple yet data-efficient solution to effectively improve fact-checking performance in dialogue. We point out unique challenges in DIALFACT such as handling the colloquialisms, coreferences and retrieval ambiguities in the error analysis to shed light on future research in this direction 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language ModelsJunyi Li, Xiaoxue Cheng, Xin Zhao, Jian-Yun Nie 等EMNLP 2023 · 被引用 224 次
- AlignScore: Evaluating Factual Consistency with A Unified Alignment FunctionYuheng Zha, Yichi Yang, Ruichen Li, Zhiting HuACL 2023 · 被引用 44 次
- InstructDial: Improving Zero and Few-shot Generalization in Dialogue through Instruction TuningPrakhar Gupta, Cathy Jiao, Yi-Ting Yeh, Shikib Mehri 等EMNLP 2022 · 被引用 26 次
- Missing Counter-Evidence Renders NLP Fact-Checking Unrealistic for MisinformationMax Glockner, Yufang Hou, Iryna GurevychEMNLP 2022 · 被引用 23 次
- WiCE: Real-World Entailment for Claims in WikipediaRyo Kamoi, Tanya Goyal, Juan Diego Rodriguez, Greg DurrettEMNLP 2023 · 被引用 21 次
它引用的顶会 Paper16
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang 等ICLR 2020 · 被引用 674 次
- Beyond Goldfish Memory: Long-Term Open-Domain ConversationJing Xu, Arthur Szlam, Jason WestonACL 2022 · 被引用 329 次
- Coreferential Reasoning Learning for Language RepresentationDeming Ye, Yankai Lin, Jiaju Du, Zhenghao Liu 等EMNLP 2020 · 被引用 164 次
相关 Paper
- CFEVER: A Chinese Fact Extraction and VERification DatasetYing-Jia Lin, Chun-Yi Lin, Chia-Jen Yeh, Yi-Ting Li 等AAAI 2024 · 被引用 8 次
- Countering Misinformation via Emotional Response GenerationDaniel Russo, Shane P. Kaszefski-Yaschuk, Jacopo Staiano, Marco GueriniEMNLP 2023 · 被引用 4 次
- DeSePtion: Dual Sequence Prediction and Adversarial Examples for Improved Fact-CheckingChristopher Hidey, Tuhin Chakrabarty, Tariq Alhindi, Siddharth Varia 等ACL 2020 · 被引用 7 次
- ViFactCheck: A New Benchmark Dataset and Methods for Multi-Domain News Fact-Checking In VietnameseTran Thai Hoa, Tran Quang Duy, Khanh Quoc Tran, Kiet Van NguyenAAAI 2025
- WatClaimCheck: A new Dataset for Claim Entailment and InferenceKashif Khan, Ruizhe Wang, Pascal PoupartACL 2022
