Claim Matching Beyond English to Scale Global Fact-Checking
Ashkan Kazemi, Kiran Garimella, Devin Gaffney, Scott Hale
摘要
Manual fact-checking does not scale well to serve the needs of the internet. This issue is further compounded in non-English contexts. In this paper, we discuss claim matching as a possible solution to scale fact-checking. We define claim matching as the task of identifying pairs of textual messages containing claims that can be served with one fact-check. We construct a novel dataset of WhatsApp tipline and public group messages alongside fact-checked claims that are first annotated for containing "claim-like statements" and then matched with potentially similar items and annotated for claim matching. Our dataset contains content in high-resource (English, Hindi) and lower-resource (Bengali, Malayalam, Tamil) languages. We train our own embedding model using knowledge distillation and a high-quality "teacher" model in order to address the imbalance in embedding quality between the low-and high-resource languages in our dataset. We provide evaluations on the performance of our solution and compare with baselines and existing state-of-the-art multilingual embedding models, namely LASER and LaBSE. We demonstrate that our performance exceeds LASER and LaBSE in all settings. We release our annotated datasets 1 , codebooks, and trained embedding model 2 to allow for further research. Table 1: Example message pairs in our data annotated for claim similarity. Item #1 Item #2 Label पािक�तान म� गनपॉइं ट पर हु ई एक डकै ती को बताया जा रहा है मु ं बई की घटना कराची पािक�तान म� घिटत लू ट को मु ं बई का बताया जा रहा है । Very Similar பாகிஸ் தானில் உள் ள இந் திய �தர் உடன�யாக ெடல் லி தி�ம் ப மத் திய அர� உத் தர� ெசய் திகள் 24/7 FLASH பாகிஸ் தானில் உள் ள இந் திய �தர் ெடல் லி தி�ம் ப மத் திய அர� உத் தர� என தகவல் ..
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Human-centered NLP Fact-checking: Co-Designing with Fact-checkers using Matchmaking for AIHoujiang Liu, Anubrata Das, Alexander Boltz, Didi Zhou 等CSCW 2024 · 被引用 23 次
- Multilingual Previously Fact-Checked Claim RetrievalMatús Pikuliak, Ivan Srba, Róbert Móro, Timo Hromadka 等EMNLP 2023 · 被引用 9 次
- ViClaim: A Multilingual Multilabel Dataset for Automatic Claim Detection in VideosPatrick Giedemann, Pius von Däniken, Jan Milan Deriu, Álvaro Rodrigo 等EMNLP 2025 · 被引用 1 次
- Multilingual vs Crosslingual Retrieval of Fact-Checked Claims: A Tale of Two ApproachesAlan Ramponi, Marco Rovera, Róbert Móro, Sara TonelliEMNLP 2025
- ProxyQA: An Alternative Framework for Evaluating Long-Form Text Generation with Large Language ModelsHaochen Tan, Zhijiang Guo, Zhan Shi, Lu Xu 等ACL 2024
它引用的顶会 Paper6
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Making Monolingual Sentence Embeddings Multilingual using Knowledge DistillationNils Reimers, Iryna GurevychEMNLP 2020 · 被引用 54 次
- That is a Known Lie: Detecting Previously Fact-Checked ClaimsShaden Shaar, Nikolay Babulkov, Giovanni Da San Martino, Preslav NakovACL 2020 · 被引用 26 次
- Where Are the Facts? Searching for Fact-checked Information to Alleviate the Spread of Fake NewsNguyen Vo, Kyumin LeeEMNLP 2020 · 被引用 4 次
相关 Paper
- Lost in Translation, Found in Spans: Identifying Claims in Multilingual Social MediaShubham Mittal, Megha Sundriyal, Preslav NakovEMNLP 2023 · 被引用 4 次
- ViFactCheck: A New Benchmark Dataset and Methods for Multi-Domain News Fact-Checking In VietnameseTran Thai Hoa, Tran Quang Duy, Khanh Quoc Tran, Kiet Van NguyenAAAI 2025
- AFaCTA: Assisting the Annotation of Factual Claim Detection with Reliable LLM AnnotatorsJingwei Ni, Minjing Shi, Dominik Stammbach, Mrinmaya Sachan 等ACL 2024
- FACTIFY-5WQA: 5W Aspect-based Fact Verification through Question AnsweringAnku Rani, S. M. Towhidul Islam Tonmoy, Dwip Dalal, Shreya Gautam 等ACL 2023 · 被引用 16 次
- Evaluating Peer Fact-Checking on WhatsAppSudhamshu Hosamane, Kriti Sharma, Tanvi Goyal, Molly Offer-Westort 等CHI 2026 · 被引用 1 次
