Efficiency and Effectiveness of LLM-Based Summarization of Evidence in Crowdsourced Fact-Checking
Kevin Roitero, Dustin Wright, Michael Soprano, Isabelle Augenstein, Stefano Mizzaro
Abstract
Evaluating the truthfulness of online content is critical for combating misinformation. This study examines the efficiency and effectiveness of crowdsourced truthfulness assessments through a comparative analysis of two approaches: one involving full-length webpages as evidence for each claim, and another using summaries for each evidence document generated with a LLM. Using an A/B testing setting, we engage a diverse pool of participants tasked with evaluating the truthfulness of statements under these conditions.
Our analysis explores both the quality of assessments and the behavioral patterns of participants. The results reveal that relying on summarized evidence offers comparable accuracy and error metrics to the standard modality while significantly improving efficiency. Workers in the Summary setting complete a significantly higher number of assessments, reducing task duration and costs. Additionally, the Summary modality maximizes internal agreement and maintains consistent reliance on and perceived usefulness of evidence, demonstrating its potential to streamline largescale truthfulness evaluations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4cea340b-b96d-4b53-8a02-25167696aaa6Builds on6
- Generating Fact Checking ExplanationsPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinACL 2020 · 130 citations
- Birds of a feather don't fact-check each other: Partisanship and the evaluation of news in Twitter's Birdwatch crowdsourced fact-checking programJennifer Allen, Cameron Martel, David G. RandCHI 2022 · 104 citations
- Human-centered NLP Fact-checking: Co-Designing with Fact-checkers using Matchmaking for AIHoujiang Liu, Anubrata Das, Alexander Boltz, Didi Zhou et al.CSCW 2024 · 23 citations
- Show Me the Work: Fact-Checkers' Requirements for Explainable Automated Fact-CheckingGreta Warren, Irina Shklovski, Isabelle AugensteinCHI 2025 · 15 citations
- Exploring Multidimensional Checkworthiness: Designing AI-assisted Claim Prioritization for Human Fact-checkersHoujiang Liu, Jacek Gwizdka, Matthew LeaseCSCW 2025 · 3 citations
Related papers
- Can The Crowd Identify Misinformation Objectively?: The Effects of Judgment Scale and Assessor's BackgroundKevin Roitero, Michael Soprano, Shaoyang Fan, Damiano Spina et al.SIGIR 2020 · 2 citations
- Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model PerformanceOmer Nahum, Nitay Calderon, Orgad Keller, Idan Szpektor et al.EMNLP 2025 · 9 citations
- The Magnitude of Truth: On Using Magnitude Estimation for Truthfulness AssessmentMichael Soprano, Denis Eduard Tapu, David La Barbera, Kevin Roitero et al.SIGIR 2025
- SAGE: A Search-AuGmented Evaluation of Large Language Models on Free-Form QASher Badshah, Ali Emami, Hassan SajjadACL 2026 · 1 citation
- The Viability of Crowdsourcing for RAG EvaluationLukas Gienapp, Tim Hagen, Maik Fröbe, Matthias Hagen et al.SIGIR 2025 · 7 citations
