Unknown Claims: Generation of Fact-Checking Training Examples from Unstructured and Structured Data
Jean-Flavien Bussotti, Luca Ragazzi, Giacomo Frisoni, Gianluca Moro, Paolo Papotti
Abstract
Computational fact-checking (FC) relies on supervised models to verify claims based on given evidence, requiring a resource-intensive process to annotate large volumes of training data. We introduce UNOWN, a novel framework that generates training instances for FC systems automatically using both textual and tabular content. UNOWN selects relevant evidence and generates supporting and refuting claims with advanced negation artifacts. Designed to be flexible, UNOWN accommodates various strategies for evidence selection and claim generation, offering unparalleled adaptability. We comprehensively evaluate UNOWN on both text-only and table+text benchmarks, including FEVEROUS, SCIFACT, and MMFC, a new multi-modal FC dataset. Our results prove that UNOWN examples are of comparable quality to expert-labeled data, even enabling models to achieve up to 5% higher accuracy. The code, data, and models are available at https: //github.com/disi-unibo-nlp/unown *
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8dedd64b-3d3e-485b-835d-e6673347b3fbCited by top-tier papers4
- ClaimDB: A Fact Verification Benchmark over Large Structured DataMichael Theologitis, Preetam Prabhu Srikar Dammu, Chirag Shah, Dan SuciuACL 2026 · 2 citations
- Beyond Static Artifacts: An Evolutionary Framework for Synthetic Claim GenerationYeqing Teng, Jiasheng Si, Shuxia Lin, Linhai Zhang et al.ACL 2026
- MPCG: Multi-Round Persona-Conditioned Generation for Modeling the Evolution of Misinformation with LLMsJun Rong Brian Chong, Yixuan Tang, Anthony Kum Hoe TungEMNLP 2025
- SQUAB: Evaluating LLM robustness to Ambiguous and Unanswerable Questions in Semantic ParsingSimone Papicchio, Luca Cagliero, Paolo PapottiEMNLP 2025
Builds on19
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang et al.ICLR 2020 · 674 citations
- Generating Training Data with Language Models: Towards Zero-Shot Language UnderstandingYu Meng, Jiaxin Huang, Yu Zhang, Jiawei HanNeurIPS 2022 · 309 citations
Related papers
- Unsupervised Pretraining for Fact Verification by Language Model DistillationAdrián Bazaga, Pietro Lio, Gos MicklemICLR 2024 · 5 citations
- Toward a Unified Framework for Unsupervised Complex Tabular ReasoningZhenyu Li, Xiuxing Li, Zhichao Duan, Bowen Dong et al.ICDE 2023 · 4 citations
- Heterogeneous Graph Reasoning for Fact Checking over Texts and TablesHaisong Gong, Weizhi Xu, Shu Wu, Qiang Liu et al.AAAI 2024 · 19 citations
- CFEVER: A Chinese Fact Extraction and VERification DatasetYing-Jia Lin, Chun-Yi Lin, Chia-Jen Yeh, Yi-Ting Li et al.AAAI 2024 · 8 citations
- Enhancing Structured Evidence Extraction for Fact VerificationZirui Wu, Nan Hu, Yansong FengEMNLP 2023 · 1 citation
