Scrutinizer: A Mixed-Initiative Approach to Large-Scale, Data-Driven Claim Verification
Georgios Karagiannis, Mohammed Saeed, Paolo Papotti, Immanuel Trummer
摘要
Organizations spend significant amounts of time and money to manually fact check text documents summarizing data. The goal of the Scrutinizer system is to reduce verification overheads by supporting human fact checkers in translating text claims into SQL queries on an database. Scrutinizer coordinates teams of human fact checkers. It reduces verification time by proposing queries or query fragments to the users. Those proposals are based on claim text classifiers, that gradually improve during the verification of a large document. In addition, Scrutinizer uses tentative execution of query candidates to narrow down the set of alternatives. The verification process is controlled by a cost-based optimizer. It optimizes the interaction with users and prioritizes claim verifications. For the latter, it considers expected verification overheads as well as the expected claim utility as training samples for the classifiers. We evaluate the Scrutinizer system using simulations and a user study with professional fact checkers, based on actual claims and data. Our experiments consistently demonstrate significant savings in verification time, without reducing result accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- DB-BERT: A Database Tuning Tool that "Reads the Manual"Immanuel TrummerSIGMOD 2022 · 被引用 71 次
- On Detecting Cherry-picked GeneralizationsYin Lin, Brit Youngmann, Yuval Moskovitch, H. V. Jagadish 等VLDB 2022 · 被引用 18 次
- Unsupervised Matching of Data and TextNaser Ahmadi, Hansjorg Sand, Paolo PapottiICDE 2022 · 被引用 18 次
- Can Large Language Models Predict Data Correlations from Column Names?Immanuel TrummerVLDB 2023 · 被引用 17 次
- Generation of Training Examples for Tabular Natural Language InferenceJean-Flavien Bussotti, Enzo Veltri, Donatello Santoro, Paolo PapottiSIGMOD 2024 · 被引用 7 次
它引用的顶会 Paper3
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang 等ICLR 2020 · 被引用 674 次
- Mining an "Anti-Knowledge Base" from Wikipedia Updates with Applications to Fact Checking and BeyondGeorgios Karagiannis, Immanuel Trummer, Saehan Jo, Shubham Khandelwal 等VLDB 2020 · 被引用 21 次
- TaPas: Weakly Supervised Table Parsing via Pre-trainingJonathan Herzig, Pawel Krzysztof Nowak, Thomas Müller, Francesco Piccinno 等ACL 2020 · 被引用 19 次
相关 Paper
- CEDAR: A System for Cost-Efficient Data-Driven Claim VerificationTharushi Jayasekara, Immanuel TrummerVLDB 2025 · 被引用 1 次
- Developing and Benchmarking Verification Algorithms to Improve Text-to-SQL GenerationTarfah Alrashed, Madhup Sukoon, David R. Karger, Natasha F. NoyVLDB 2026
- SQL-Checker: Error Detection and Labeling for Text-to-SQL with Interpretability AnalysisXingyu Ma, Xin Tian, Lingxiang Wu, Xuepeng Wang 等WWW 2026
- Factoring Fact-Checks: Structured Information Extraction from Fact-Checking ArticlesShan Jiang, Simon Baumgartner, Abe Ittycheriah, Cong YuWWW 2020 · 被引用 28 次
- SpotIt: Evaluating Text-to-SQL Evaluation with Formal VerificationRocky Klopfenstein, Yang He, Andrew Tremante, Yuepeng Wang 等ICLR 2026 · 被引用 6 次
