CEDAR: A System for Cost-Efficient Data-Driven Claim Verification
Tharushi Jayasekara, Immanuel Trummer
Abstract
We present CEDAR, a system for cost-efficient, data-driven claim verification. CEDAR takes as input a collection of text documents, containing claims that can be verified from relational data. The system uses large language models (LLMs) to map claims to SQL queries that can be used for claim verification. While LLMs like GPT-4 are nowadays able to map claims to queries with high accuracy, using them is expensive. This is why CEDAR implements multiple verification approaches, ranging from zero-shot LLM invocations to iterative, agent-based approaches, that realize different tradeoffs between accuracy and costs. The system may apply multiple methods to the same claim, starting with cheaper methods and resorting to more expensive versions in case of failures. CEDAR uses cost-based optimization to derive an optimal order of verification methods and an optimal number of re-tries (with randomization) for each method, enabling users to trade costs for accuracy via tuning parameters. The experiments on real data, including newspaper and Wikipedia articles, show that CEDAR achieves significantly higher accuracy than prior methods for data-driven fact-checking.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang et al.ICLR 2020 · 674 citations
- Text-to-SQL Empowered by Large Language Models: A Benchmark EvaluationDawei Gao, Haibin Wang, Yaliang Li, Xiuyu Sun et al.VLDB 2024 · 609 citations
- TAPEX: Table Pre-training via Learning a Neural SQL ExecutorQian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi et al.ICLR 2022 · 347 citations
- Mining an "Anti-Knowledge Base" from Wikipedia Updates with Applications to Fact Checking and BeyondGeorgios Karagiannis, Immanuel Trummer, Saehan Jo, Shubham Khandelwal et al.VLDB 2020 · 21 citations
Related papers
- Long-form factuality in large language modelsJerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu et al.NeurIPS 2024 · 182 citations
- Retrieval-Based Prompt Selection for Code-Related Few-Shot LearningNoor Nashid, Mifta Sintaha, Ali MesbahICSE 2023 · 156 citations
- ClaimDB: A Fact Verification Benchmark over Large Structured DataMichael Theologitis, Preetam Prabhu Srikar Dammu, Chirag Shah, Dan SuciuACL 2026 · 2 citations
- Scrutinizer: A Mixed-Initiative Approach to Large-Scale, Data-Driven Claim VerificationGeorgios Karagiannis, Mohammed Saeed, Paolo Papotti, Immanuel TrummerVLDB 2020
- FinDVer: Explainable Claim Verification over Long and Hybrid-content Financial DocumentsYilun Zhao, Yitao Long, Tintin Jiang, Chengye Wang et al.EMNLP 2024 · 3 citations
