"The Data Says Otherwise" - Towards Automated Fact-checking and Communication of Data Claims
Yu Fu, Shunan Guo, Jane Hoffswell, Victor S. Bursztyn, Ryan A. Rossi, John T. Stasko
Abstract
Fact-checking data claims requires data evidence retrieval and analysis, which can become tedious and intractable when done manually. This work presents Aletheia, an automated fact-checking prototype designed to facilitate data claims verification and enhance data evidence communication. For verification, we utilize a pre-trained LLM to parse the semantics for evidence retrieval. To effectively communicate the data evidence, we design representations in two forms: data tables and visualizations, tailored to various data fact types. Additionally, we design interactions that showcase a real-world application of these techniques. We evaluate the performance of two core NLP tasks with a curated dataset comprising 400 data claims and compare the two representation forms regarding viewers’ assessment time, confidence, and preference via a user study with 20 participants. The evaluation offers insights into the feasibility and bottlenecks of using LLMs for data fact-checking tasks, potential advantages and disadvantages of using visualizations over data tables, and design recommendations for presenting data evidence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Jupybara: Operationalizing a Design Space for Actionable Data Analysis and Storytelling with LLMsHuichen Will Wang, Larry Birnbaum, Vidya SetlurCHI 2025 · 11 citations
- RAGTrace: Understanding and Refining Retrieval-Generation Dynamics in Retrieval-Augmented GenerationSizhe Cheng, Jiaping Li, Huanchen Wang, Yuxin MaUIST 2025 · 6 citations
- Behavioral Indicators of Overreliance During Interaction with Conversational Language ModelsChang Liu, Qinyi Zhou, Xinjie Shen, Xingyu Bruce Liu et al.CHI 2026 · 4 citations
- TableTale: Reviving the Narrative Interplay Between Data Tables and Text in Scientific PapersLiangwei Wang, Zhengxuan Zhang, Yifan Cao, Fugee Tsung et al.CHI 2026 · 2 citations
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan et al.VLDB 2021 · 484 citations
- AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model PromptsTongshuang Wu, Michael Terry, Carrie Jun CaiCHI 2022 · 465 citations
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
- Calliope: Automatic Visual Data Story Generation from a SpreadsheetDanqing Shi, Xinyue Xu, Fuling Sun, Yang Shi et al.IEEE VIS 2020 · 179 citations
Related papers
- Reasoning Over Semantic-Level Graph for Fact CheckingWanjun Zhong, Jingjing Xu, Duyu Tang, Zenan Xu et al.ACL 2020 · 154 citations
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang et al.ICLR 2020 · 674 citations
- A Fact-Checking Framework with Denoising Evidence Retrieval and LLM-Based Debate VerificationJun Yang, Yuhan Bai, Dandan Song, Zhijing Wu et al.WWW 2026
- ClaimDB: A Fact Verification Benchmark over Large Structured DataMichael Theologitis, Preetam Prabhu Srikar Dammu, Chirag Shah, Dan SuciuACL 2026 · 2 citations
- Heterogeneous Graph Reasoning for Fact Checking over Texts and TablesHaisong Gong, Weizhi Xu, Shu Wu, Qiang Liu et al.AAAI 2024 · 19 citations
