VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking
Mark Rothermel, Marcus Kornmann, Marcus Rohrbach, Anna Rohrbach
摘要
The growing scale of online misinformation urgently demands Automated Fact-Checking (AFC). Existing benchmarks for evaluating AFC systems, however, are largely limited in terms of task scope, modalities, domain, language diversity, realism, or coverage of misinformation types. Critically, they are static, thus subject to data leakage as their claims enter the pretraining corpora of LLMs. As a result, benchmark performance no longer reliably reflects the actual ability to verify claims. We introduce Verified Theses and Statements (VeriTaS), the first dynamic benchmark for multimodal AFC, designed to remain robust under ongoing large-scale pretraining of foundation models. VeriTaS currently comprises 25,000 real-world claims from 104 professional fact-checking organizations across 54 languages, covering textual and audiovisual content. Claims are added quarterly via a fully automated seven-stage pipeline that normalizes claim formulation, retrieves original media, and maps heterogeneous expert verdicts to a novel, standardized, and disentangled scoring scheme with textual justifications. Through human evaluation, we demonstrate that the automated annotations closely match human judgments. We commit to updating VeriTaS in the future, establishing a leakage-resistant benchmark, supporting meaningful AFC evaluation in the era of rapidly evolving foundation models. The code and data are publicly available under https://veritas.mai.informatik.tu-darmstadt.de .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- NewsCLIPpings: Automatic Generation of Out-of-Context Multimodal MediaGrace Luo, Trevor Darrell, Anna RohrbachEMNLP 2021 · 被引用 58 次
- Missing Counter-Evidence Renders NLP Fact-Checking Unrealistic for MisinformationMax Glockner, Yufang Hou, Iryna GurevychEMNLP 2022 · 被引用 23 次
- "Image, Tell me your story!" Predicting the original meta-context of visual misinformationJonathan Tonglet, Marie-Francine Moens, Iryna GurevychEMNLP 2024 · 被引用 6 次
- HintsOfTruth: A Multimodal Checkworthiness Detection Dataset with Real and Synthetic ClaimsMichiel van der Meer, Pavel Korshunov, Sébastien Marcel, Lonneke van der PlasACL 2025 · 被引用 5 次
- FACTIFY3M: A benchmark for multimodal fact verification with explainability through 5W Question-AnsweringMegha Chakraborty, Khushbu Pahwa, Anku Rani, Shreyas Chatterjee 等EMNLP 2023 · 被引用 3 次
相关 Paper
- DEFAME: Dynamic Evidence-based FAct-checking with Multimodal ExpertsTobias Braun, Mark Rothermel, Marcus Rohrbach, Anna RohrbachICML 2025
- Automated Justification Production for Claim Veracity in Fact Checking: A Survey on Architectures and ApproachesIslam Eldifrawi, Shengrui Wang, Amine TrabelsiACL 2024 · 被引用 4 次
- LiveFact: A Dynamic, Time-Aware Benchmark for LLM-Driven Fake News DetectionCheng Xu, Changhong Jin, Yingjie Niu, Nan Yan 等ACL 2026
- INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMsJunqi Yang, Yuecong Min, Jie Zhang, Shiguang Shan 等ACL 2026 · 被引用 7 次
- FakeWorld 1.0: An Omni-modal Benchmark for Fake Media and ContentYifeng Gao, Yifan Ding, Li Wang, Feida Huang 等ICML 2026
