DialFact: A Benchmark for Fact-Checking in Dialogue
Prakhar Gupta, Chien-Sheng Wu, Wenhao Liu, Caiming Xiong
Abstract
Fact-checking is an essential tool to mitigate the spread of misinformation and disinformation. We introduce the task of fact-checking in dialogue, which is a relatively unexplored area. We construct DIALFACT, a testing benchmark dataset of 22,245 annotated conversational claims, paired with pieces of evidence from Wikipedia. There are three sub-tasks in DIALFACT: 1) Verifiable claim detection task distinguishes whether a response carries verifiable factual information; 2) Evidence retrieval task retrieves the most relevant Wikipedia snippets as evidence; 3) Claim verification task predicts a dialogue response to be supported, refuted, or not enough information. We found that existing fact-checking models trained on non-dialogue data like FEVER (Thorne et al., 2018) fail to perform well on our task, and thus, we propose a simple yet data-efficient solution to effectively improve fact-checking performance in dialogue. We point out unique challenges in DIALFACT such as handling the colloquialisms, coreferences and retrieval ambiguities in the error analysis to shed light on future research in this direction 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ab5f270-8998-486f-806e-3edcfeff67d1Cited by top-tier papers21
- HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language ModelsJunyi Li, Xiaoxue Cheng, Xin Zhao, Jian-Yun Nie et al.EMNLP 2023 · 224 citations
- AlignScore: Evaluating Factual Consistency with A Unified Alignment FunctionYuheng Zha, Yichi Yang, Ruichen Li, Zhiting HuACL 2023 · 44 citations
- InstructDial: Improving Zero and Few-shot Generalization in Dialogue through Instruction TuningPrakhar Gupta, Cathy Jiao, Yi-Ting Yeh, Shikib Mehri et al.EMNLP 2022 · 26 citations
- Missing Counter-Evidence Renders NLP Fact-Checking Unrealistic for MisinformationMax Glockner, Yufang Hou, Iryna GurevychEMNLP 2022 · 23 citations
- WiCE: Real-World Entailment for Claims in WikipediaRyo Kamoi, Tanya Goyal, Juan Diego Rodriguez, Greg DurrettEMNLP 2023 · 21 citations
Builds on16
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang et al.ICLR 2020 · 674 citations
- Beyond Goldfish Memory: Long-Term Open-Domain ConversationJing Xu, Arthur Szlam, Jason WestonACL 2022 · 329 citations
- Coreferential Reasoning Learning for Language RepresentationDeming Ye, Yankai Lin, Jiaju Du, Zhenghao Liu et al.EMNLP 2020 · 164 citations
Related papers
- CFEVER: A Chinese Fact Extraction and VERification DatasetYing-Jia Lin, Chun-Yi Lin, Chia-Jen Yeh, Yi-Ting Li et al.AAAI 2024 · 8 citations
- Countering Misinformation via Emotional Response GenerationDaniel Russo, Shane P. Kaszefski-Yaschuk, Jacopo Staiano, Marco GueriniEMNLP 2023 · 4 citations
- DeSePtion: Dual Sequence Prediction and Adversarial Examples for Improved Fact-CheckingChristopher Hidey, Tuhin Chakrabarty, Tariq Alhindi, Siddharth Varia et al.ACL 2020 · 7 citations
- ViFactCheck: A New Benchmark Dataset and Methods for Multi-Domain News Fact-Checking In VietnameseTran Thai Hoa, Tran Quang Duy, Khanh Quoc Tran, Kiet Van NguyenAAAI 2025
- WatClaimCheck: A new Dataset for Claim Entailment and InferenceKashif Khan, Ruizhe Wang, Pascal PoupartACL 2022
