Document-level Claim Extraction and Decontextualisation for Fact-Checking
Zhenyun Deng, Michael Sejr Schlichtkrull, Andreas Vlachos
Abstract
Selecting which claims to check is a timeconsuming task for human fact-checkers, especially from documents consisting of multiple sentences and containing multiple claims. However, existing claim extraction approaches focus more on identifying and extracting claims from individual sentences, e.g., identifying whether a sentence contains a claim or the exact boundaries of the claim within a sentence. In this paper, we propose a method for documentlevel claim extraction for fact-checking, which aims to extract check-worthy claims from documents and decontextualise them so that they can be understood out of context. Specifically, we first recast claim extraction as extractive summarization in order to identify central sentences from documents, then rewrite them to include necessary context from the originating document through sentence decontextualisation. Evaluation with both automatic metrics and a fact-checking professional shows that our method is able to extract check-worthy claims from documents more accurately than previous work, while also improving evidence retrieval.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- When Misinformation Speaks and Converses: Rethinking Fact-Checking in Audio PlatformsChaewan Chun, Delvin Ce Zhang, Dongwon LeeACL 2026
- Improving Zero-shot Sentence Decontextualisation with Content Selection and PlanningZhenyun Deng, Yulong Chen, Andreas VlachosEMNLP 2025
- TracSum: A New Benchmark for Aspect-Based Summarization with Sentence-Level Traceability in Medical DomainBohao Chu, Meijie Li, Sameh Frihat, Chengyu Gu et al.EMNLP 2025
- The Psychology of Falsehood: A Human-Centric Survey of Misinformation DetectionArghodeep Nandi, Megha Sundriyal, Euna Mehnaz Khan, Jikai Sun et al.EMNLP 2025
Builds on3
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Empowering the Fact-checkers! Automatic Identification of Claim Spans on TwitterMegha Sundriyal, Atharva Kulkarni, Vaibhav Pulastya, Md. Shad Akhtar et al.EMNLP 2022 · 12 citations
Related papers
- Measuring and Enhancing Human Value Alignment in Zero-Shot Document-Level Claim ExtractionYuanzhen Hao, Desheng WuWWW 2026
- Towards Effective Extraction and Evaluation of Factual ClaimsDasha Metropolitansky, Jonathan LarsonACL 2025 · 17 citations
- Is the Top Still Spinning? Evaluating Subjectivity in Narrative UnderstandingMelanie Subbiah, Akankshya Mishra, Grace Kim, Liyan Tang et al.EMNLP 2025
- Varifocal Question Generation for Fact-checkingNedjma Ousidhoum, Zhangdie Yuan, Andreas VlachosEMNLP 2022 · 11 citations
- DnDScore: Decontextualization and Decomposition for Factuality Verification in Long-Form Text GenerationMiriam Wanner, Benjamin Van Durme, Mark DredzeEMNLP 2025 · 17 citations
