CHECKWHY: Causal Fact Verification via Argument Structure
Jiasheng Si, Yibo Zhao, Yingjie Zhu, Haiyang Zhu, Wenpeng Lu, Deyu Zhou
摘要
With the growing complexity of fact verification tasks, the concern with "thoughtful" reasoning capabilities is increasing. However, recent fact verification benchmarks mainly focus on checking a narrow scope of semantic factoids within claims and lack an explicit logical reasoning process. In this paper, we introduce CHECKWHY, a challenging dataset tailored to a novel causal fact verification task: checking the truthfulness of the causal relation within claims through rigorous reasoning steps. CHECKWHY consists of over 19K "why" claimevidence-argument structure triplets with supports, refutes, and not enough info labels. Each argument structure is composed of connected evidence, representing the reasoning process that begins with foundational evidence and progresses toward claim establishment. Through extensive experiments on state-of-the-art models, we validate the importance of incorporating the argument structure for causal fact verification. Moreover, the automated and human evaluation of argument structure generation reveals the difficulty in producing satisfying argument structure by fine-tuned models or Chainof-Thought prompted LLMs, leaving considerable room for future improvements 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and ReasoningYingjie Zhu, Xuefeng Bai, Kehai Chen, Yang Xiang 等ACL 2025 · 被引用 15 次
- The Missing Parts: Augmenting Fact Verification with Half Truth DetectionYixuan Tang, Jincheng Wang, Anthony Kum Hoe TungEMNLP 2025 · 被引用 6 次
- Coordinating Search-Informed Reasoning and Reasoning-Guided Search in Claim VerificationQisheng Hu, Quanyu Long, Wenya WangACL 2026 · 被引用 4 次
- Long-Form Information Alignment Evaluation Beyond Atomic FactsDanna Zheng, Mirella Lapata, Jeff Z. PanEMNLP 2025 · 被引用 2 次
- METER: Evaluating Multi-Level Contextual Causal Reasoning in Large Language ModelsPengfeng Li, Chen Huang, Chaoqun Hao, Hongyao Chen 等ACL 2026 · 被引用 1 次
它引用的顶会 Paper22
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- CodeT5+: Open Code Large Language Models for Code Understanding and GenerationYue Wang, Hung Le, Akhilesh Gotmare, Nghi D. Q. Bui 等EMNLP 2023 · 被引用 339 次
相关 Paper
- Causal Walk: Debiasing Multi-Hop Fact Verification with Front-Door AdjustmentCongzhi Zhang, Linhai Zhang, Deyu ZhouAAAI 2024 · 被引用 32 次
- A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning ChainsAlon Jacovi, Yonatan Bitton, Bernd Bohnet, Jonathan Herzig 等ACL 2024 · 被引用 8 次
- CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question AnsweringMingfang Zhang, Jingjing Pan, Ashutosh Kumar, Rajat Saini 等CVPR 2026 · 被引用 1 次
- WikiWhy: Answering and Explaining Cause-and-Effect QuestionsMatthew Ho, Aditya Sharma, Justin Chang, Michael Saxon 等ICLR 2023 · 被引用 8 次
- e-CARE: a New Dataset for Exploring Explainable Causal ReasoningLi Du, Xiao Ding, Kai Xiong, Ting Liu 等ACL 2022
