Efficiency and Effectiveness of LLM-Based Summarization of Evidence in Crowdsourced Fact-Checking
Kevin Roitero, Dustin Wright, Michael Soprano, Isabelle Augenstein, Stefano Mizzaro
摘要
Evaluating the truthfulness of online content is critical for combating misinformation. This study examines the efficiency and effectiveness of crowdsourced truthfulness assessments through a comparative analysis of two approaches: one involving full-length webpages as evidence for each claim, and another using summaries for each evidence document generated with a LLM. Using an A/B testing setting, we engage a diverse pool of participants tasked with evaluating the truthfulness of statements under these conditions.
Our analysis explores both the quality of assessments and the behavioral patterns of participants. The results reveal that relying on summarized evidence offers comparable accuracy and error metrics to the standard modality while significantly improving efficiency. Workers in the Summary setting complete a significantly higher number of assessments, reducing task duration and costs. Additionally, the Summary modality maximizes internal agreement and maintains consistent reliance on and perceived usefulness of evidence, demonstrating its potential to streamline largescale truthfulness evaluations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Generating Fact Checking ExplanationsPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinACL 2020 · 被引用 130 次
- Birds of a feather don't fact-check each other: Partisanship and the evaluation of news in Twitter's Birdwatch crowdsourced fact-checking programJennifer Allen, Cameron Martel, David G. RandCHI 2022 · 被引用 104 次
- Human-centered NLP Fact-checking: Co-Designing with Fact-checkers using Matchmaking for AIHoujiang Liu, Anubrata Das, Alexander Boltz, Didi Zhou 等CSCW 2024 · 被引用 23 次
- Show Me the Work: Fact-Checkers' Requirements for Explainable Automated Fact-CheckingGreta Warren, Irina Shklovski, Isabelle AugensteinCHI 2025 · 被引用 15 次
- Exploring Multidimensional Checkworthiness: Designing AI-assisted Claim Prioritization for Human Fact-checkersHoujiang Liu, Jacek Gwizdka, Matthew LeaseCSCW 2025 · 被引用 3 次
相关 Paper
- Can The Crowd Identify Misinformation Objectively?: The Effects of Judgment Scale and Assessor's BackgroundKevin Roitero, Michael Soprano, Shaoyang Fan, Damiano Spina 等SIGIR 2020 · 被引用 2 次
- Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model PerformanceOmer Nahum, Nitay Calderon, Orgad Keller, Idan Szpektor 等EMNLP 2025 · 被引用 9 次
- The Magnitude of Truth: On Using Magnitude Estimation for Truthfulness AssessmentMichael Soprano, Denis Eduard Tapu, David La Barbera, Kevin Roitero 等SIGIR 2025
- SAGE: A Search-AuGmented Evaluation of Large Language Models on Free-Form QASher Badshah, Ali Emami, Hassan SajjadACL 2026 · 被引用 1 次
- The Viability of Crowdsourcing for RAG EvaluationLukas Gienapp, Tim Hagen, Maik Fröbe, Matthias Hagen 等SIGIR 2025 · 被引用 7 次
