Unknown Claims: Generation of Fact-Checking Training Examples from Unstructured and Structured Data
Jean-Flavien Bussotti, Luca Ragazzi, Giacomo Frisoni, Gianluca Moro, Paolo Papotti
摘要
Computational fact-checking (FC) relies on supervised models to verify claims based on given evidence, requiring a resource-intensive process to annotate large volumes of training data. We introduce UNOWN, a novel framework that generates training instances for FC systems automatically using both textual and tabular content. UNOWN selects relevant evidence and generates supporting and refuting claims with advanced negation artifacts. Designed to be flexible, UNOWN accommodates various strategies for evidence selection and claim generation, offering unparalleled adaptability. We comprehensively evaluate UNOWN on both text-only and table+text benchmarks, including FEVEROUS, SCIFACT, and MMFC, a new multi-modal FC dataset. Our results prove that UNOWN examples are of comparable quality to expert-labeled data, even enabling models to achieve up to 5% higher accuracy. The code, data, and models are available at https: //github.com/disi-unibo-nlp/unown *
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- ClaimDB: A Fact Verification Benchmark over Large Structured DataMichael Theologitis, Preetam Prabhu Srikar Dammu, Chirag Shah, Dan SuciuACL 2026 · 被引用 2 次
- Beyond Static Artifacts: An Evolutionary Framework for Synthetic Claim GenerationYeqing Teng, Jiasheng Si, Shuxia Lin, Linhai Zhang 等ACL 2026
- MPCG: Multi-Round Persona-Conditioned Generation for Modeling the Evolution of Misinformation with LLMsJun Rong Brian Chong, Yixuan Tang, Anthony Kum Hoe TungEMNLP 2025
- SQUAB: Evaluating LLM robustness to Ambiguous and Unanswerable Questions in Semantic ParsingSimone Papicchio, Luca Cagliero, Paolo PapottiEMNLP 2025
它引用的顶会 Paper19
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang 等ICLR 2020 · 被引用 674 次
- Generating Training Data with Language Models: Towards Zero-Shot Language UnderstandingYu Meng, Jiaxin Huang, Yu Zhang, Jiawei HanNeurIPS 2022 · 被引用 309 次
相关 Paper
- Unsupervised Pretraining for Fact Verification by Language Model DistillationAdrián Bazaga, Pietro Lio, Gos MicklemICLR 2024 · 被引用 5 次
- Toward a Unified Framework for Unsupervised Complex Tabular ReasoningZhenyu Li, Xiuxing Li, Zhichao Duan, Bowen Dong 等ICDE 2023 · 被引用 4 次
- Heterogeneous Graph Reasoning for Fact Checking over Texts and TablesHaisong Gong, Weizhi Xu, Shu Wu, Qiang Liu 等AAAI 2024 · 被引用 19 次
- CFEVER: A Chinese Fact Extraction and VERification DatasetYing-Jia Lin, Chun-Yi Lin, Chia-Jen Yeh, Yi-Ting Li 等AAAI 2024 · 被引用 8 次
- Enhancing Structured Evidence Extraction for Fact VerificationZirui Wu, Nan Hu, Yansong FengEMNLP 2023 · 被引用 1 次
