Scrutinizer: A Mixed-Initiative Approach to Large-Scale, Data-Driven Claim Verification
Georgios Karagiannis, Mohammed Saeed, Paolo Papotti, Immanuel Trummer
Abstract
Organizations spend significant amounts of time and money to manually fact check text documents summarizing data. The goal of the Scrutinizer system is to reduce verification overheads by supporting human fact checkers in translating text claims into SQL queries on an database. Scrutinizer coordinates teams of human fact checkers. It reduces verification time by proposing queries or query fragments to the users. Those proposals are based on claim text classifiers, that gradually improve during the verification of a large document. In addition, Scrutinizer uses tentative execution of query candidates to narrow down the set of alternatives. The verification process is controlled by a cost-based optimizer. It optimizes the interaction with users and prioritizes claim verifications. For the latter, it considers expected verification overheads as well as the expected claim utility as training samples for the classifiers. We evaluate the Scrutinizer system using simulations and a user study with professional fact checkers, based on actual claims and data. Our experiments consistently demonstrate significant savings in verification time, without reducing result accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f69f466-7494-4377-a98e-0bbc4db3c94cCited by top-tier papers7
- DB-BERT: A Database Tuning Tool that "Reads the Manual"Immanuel TrummerSIGMOD 2022 · 71 citations
- On Detecting Cherry-picked GeneralizationsYin Lin, Brit Youngmann, Yuval Moskovitch, H. V. Jagadish et al.VLDB 2022 · 18 citations
- Unsupervised Matching of Data and TextNaser Ahmadi, Hansjorg Sand, Paolo PapottiICDE 2022 · 18 citations
- Can Large Language Models Predict Data Correlations from Column Names?Immanuel TrummerVLDB 2023 · 17 citations
- Generation of Training Examples for Tabular Natural Language InferenceJean-Flavien Bussotti, Enzo Veltri, Donatello Santoro, Paolo PapottiSIGMOD 2024 · 7 citations
Builds on3
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang et al.ICLR 2020 · 674 citations
- Mining an "Anti-Knowledge Base" from Wikipedia Updates with Applications to Fact Checking and BeyondGeorgios Karagiannis, Immanuel Trummer, Saehan Jo, Shubham Khandelwal et al.VLDB 2020 · 21 citations
- TaPas: Weakly Supervised Table Parsing via Pre-trainingJonathan Herzig, Pawel Krzysztof Nowak, Thomas Müller, Francesco Piccinno et al.ACL 2020 · 19 citations
Related papers
- CEDAR: A System for Cost-Efficient Data-Driven Claim VerificationTharushi Jayasekara, Immanuel TrummerVLDB 2025 · 1 citation
- Developing and Benchmarking Verification Algorithms to Improve Text-to-SQL GenerationTarfah Alrashed, Madhup Sukoon, David R. Karger, Natasha F. NoyVLDB 2026
- SQL-Checker: Error Detection and Labeling for Text-to-SQL with Interpretability AnalysisXingyu Ma, Xin Tian, Lingxiang Wu, Xuepeng Wang et al.WWW 2026
- Factoring Fact-Checks: Structured Information Extraction from Fact-Checking ArticlesShan Jiang, Simon Baumgartner, Abe Ittycheriah, Cong YuWWW 2020 · 28 citations
- SpotIt: Evaluating Text-to-SQL Evaluation with Formal VerificationRocky Klopfenstein, Yang He, Andrew Tremante, Yuepeng Wang et al.ICLR 2026 · 6 citations
