PASTA: Table-Operations Aware Fact Verification via Sentence-Table Cloze Pre-training
Zihui Gu, Ju Fan, Nan Tang, Preslav Nakov, Xiaoman Zhao, Xiaoyong Du
摘要
Fact verification has attracted a lot of research attention recently, e.g., in journalism, marketing, and policymaking, as misinformation and disinformation online can sway one's opinion and affect one's actions. While fact-checking is a hard task in general, in many cases, false statements can be easily debunked based on analytics over tables with reliable information. Hence, table-based fact verification has recently emerged as an important and growing research area. Yet, progress has been limited due to the lack of datasets that can be used to pre-train language models (LMs) to be aware of common table operations, such as aggregating a column or comparing tuples. To bridge this gap, in this paper we introduce PASTA, a novel state-of-the-art framework for table-based fact verification via pre-training with synthesized sentence-table cloze questions. In particular, we design six types of common sentence-table cloze tasks, including Filter, Aggregation, Superlative, Comparative, Ordinal, and Unique, based on which we synthesize a large corpus consisting of 1.2 million sentence-table pairs from WikiTables. PASTA uses a recent pre-trained LM, DeBERTaV3, and further pretrains it on our corpus. Our experimental results show that PASTA achieves new state-ofthe-art performance on two table-based fact verification benchmarks: TabFact and SEM-TAB-FACTS. In particular, on the complex set of TabFact, which contains multiple operations, PASTA largely outperforms the previous state of the art by 4.7 points (85.6% vs. 80.9%), and the gap between PASTA and human performance on the small TabFact test set is narrowed to just 1.5 points (90.6% vs. 92.1%).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Chain-of-Table: Evolving Tables in the Reasoning Chain for Table UnderstandingZilong Wang, Hao Zhang, Chun-Liang Li, Julian Martin Eisenschlos 等ICLR 2024 · 被引用 244 次
- Large Language Models are Versatile Decomposers: Decomposing Evidence and Questions for Table-based ReasoningYunhu Ye, Binyuan Hui, Min Yang, Binhua Li 等SIGIR 2023 · 被引用 75 次
- Combining Small Language Models and Large Language Models for Zero-Shot NL2SQLJu Fan, Zihui Gu, Songyue Zhang, Yuxin Zhang 等VLDB 2024 · 被引用 71 次
- Unicorn: A Unified Multi-tasking Model for Supporting Matching Tasks in Data IntegrationJianhong Tu, Ju Fan, Nan Tang, Peng Wang 等SIGMOD 2023 · 被引用 34 次
- CABINET: Content Relevance-based Noise Reduction for Table Question AnsweringSohan Patnaik, Heril Changwal, Milan Aggarwal, Sumit Bhatia 等ICLR 2024 · 被引用 34 次
它引用的顶会 Paper10
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang 等ICLR 2020 · 被引用 674 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 被引用 417 次
- TAPEX: Table Pre-training via Learning a Neural SQL ExecutorQian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi 等ICLR 2022 · 被引用 347 次
- LogicalFactChecker: Leveraging Logical Operations for Fact Checking with Graph Module NetworkWanjun Zhong, Duyu Tang, Zhangyin Feng, Nan Duan 等ACL 2020 · 被引用 51 次
相关 Paper
- Program Enhanced Fact Verification with Verbalization and Graph Attention NetworkXiaoyu Yang, Feng Nie, Yufei Feng, Quan Liu 等EMNLP 2020 · 被引用 48 次
- ReasTAP: Injecting Table Reasoning Skills During Pre-training via Synthetic Reasoning ExamplesYilun Zhao, Linyong Nan, Zhenting Qi, Rui Zhang 等EMNLP 2022 · 被引用 15 次
- A Multi-Task Learning Framework for Reading Comprehension of Scientific Tabular DataXu Yang, Meihui Zhang, Ju Fan, Zeyu Luo 等ICDE 2024 · 被引用 1 次
- RePanda: Pandas-powered Tabular Verification and ReasoningAtoosa Malemir Chegini, Keivan Rezaei, Hamid Eghbalzadeh, Soheil FeiziACL 2025
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
