TabFact: A Large-scale Dataset for Table-based Fact Verification
Wenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang, Hong Wang, Shiyang Li, Xiyou Zhou, William Yang Wang
Abstract
The problem of verifying whether a textual hypothesis holds based on the given evidence, also known as fact verification, plays an important role in the study of natural language understanding and semantic representation. However, existing studies are mainly restricted to dealing with unstructured evidence (e.g., natural language sentences and documents, news, etc), while verification under structured evidence, such as tables, graphs, and databases, remains under-explored. This paper specifically aims to study the fact verification given semi-structured data as evidence. To this end, we construct a large-scale dataset called TabFact with 16k Wikipedia tables as the evidence for 118k human-annotated natural language statements, which are labeled as either ENTAILED or REFUTED. TabFact is challenging since it involves both soft linguistic reasoning and hard symbolic reasoning. To address these reasoning challenges, we design two different models: Table-BERT and Latent Program Algorithm (LPA). Table-BERT leverages the state-of-the-art pre-trained language model to encode the linearized tables and statements into continuous vectors for verification. LPA parses statements into programs and executes them against the tables to obtain the returned binary value for verification. Both methods achieve similar accuracy but still lag far behind human performance. We also perform a comprehensive analysis to demonstrate great future opportunities. The data and code of the dataset are provided in this https URL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 01cca415-0ea9-47a7-b5f6-07f245fc62cbCited by top-tier papers188
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
- TAPEX: Table Pre-training via Learning a Neural SQL ExecutorQian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi et al.ICLR 2022 · 347 citations
- Chain-of-Table: Evolving Tables in the Reasoning Chain for Table UnderstandingZilong Wang, Hao Zhang, Chun-Liang Li, Julian Martin Eisenschlos et al.ICLR 2024 · 244 citations
- MultiModalQA: complex question answering over text, tables and imagesAlon Talmor, Ori Yoran, Amnon Catav, Dan Lahav et al.ICLR 2021 · 229 citations
- MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training RecipeTianyu Yu, Zefan Wang, Chongyi Wang, Fuwei Huang et al.CVPR 2026 · 179 citations
Related papers
- Learning to Generate Programs for Table Fact Verification via Structure-Aware Semantic ParsingSuixin Ou, Yongmei LiuACL 2022 · 14 citations
- Program Enhanced Fact Verification with Verbalization and Graph Attention NetworkXiaoyu Yang, Feng Nie, Yufei Feng, Quan Liu et al.EMNLP 2020 · 48 citations
- PASTA: Table-Operations Aware Fact Verification via Sentence-Table Cloze Pre-trainingZihui Gu, Ju Fan, Nan Tang, Preslav Nakov et al.EMNLP 2022 · 19 citations
- LogicalFactChecker: Leveraging Logical Operations for Fact Checking with Graph Module NetworkWanjun Zhong, Duyu Tang, Zhangyin Feng, Nan Duan et al.ACL 2020 · 51 citations
- Logic-level Evidence Retrieval and Graph-based Verification Network for Table-based Fact VerificationQi Shi, Yu Zhang, Qingyu Yin, Ting LiuEMNLP 2021
