Multilingual vs Crosslingual Retrieval of Fact-Checked Claims: A Tale of Two Approaches
Alan Ramponi, Marco Rovera, Róbert Móro, Sara Tonelli
Abstract
Retrieval of previously fact-checked claims is a well-established task, whose automation can assist professional fact-checkers in the initial steps of information verification. Previous works have mostly tackled the task monolingually, i.e., having both the input and the retrieved claims in the same language. However, especially for languages with a limited availability of fact-checks and in case of global narratives, such as pandemics, wars, or international politics, it is crucial to be able to retrieve claims across languages. In this work, we examine strategies to improve the multilingual and crosslingual performance, namely selection of negative examples (in the supervised) and re-ranking (in the unsupervised setting). We evaluate all approaches on a dataset containing posts and claims in 47 languages (283 language combinations). We observe that the best results are obtained by using LLMbased re-ranking, followed by fine-tuning with negative examples sampled using a sentence similarity-based strategy. Most importantly, we show that crosslinguality is a setup with its own unique characteristics compared to the multilingual setup. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 205aadea-8ce2-4f99-8abe-f472894a661dBuilds on6
- Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking AgentsWeiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang et al.EMNLP 2023 · 182 citations
- That is a Known Lie: Detecting Previously Fact-Checked ClaimsShaden Shaar, Nikolay Babulkov, Giovanni Da San Martino, Preslav NakovACL 2020 · 26 citations
- MMTEB: Massive Multilingual Text Embedding BenchmarkKenneth C. Enevoldsen, Isaac Chung, Imene Kerboua, Márton Kardos et al.ICLR 2025 · 10 citations
- Multilingual Previously Fact-Checked Claim RetrievalMatús Pikuliak, Ivan Srba, Róbert Móro, Timo Hromadka et al.EMNLP 2023 · 9 citations
- Claim Matching Beyond English to Scale Global Fact-CheckingAshkan Kazemi, Kiran Garimella, Devin Gaffney, Scott HaleACL 2021
Related papers
- Lost in Translation, Found in Spans: Identifying Claims in Multilingual Social MediaShubham Mittal, Megha Sundriyal, Preslav NakovEMNLP 2023 · 4 citations
- LexCLiPR: Cross-Lingual Paragraph Retrieval from Legal JudgmentsRohit Upadhya, T. Y. S. S. SantoshACL 2025 · 3 citations
- NewsClaims: A New Benchmark for Claim Detection from News with Attribute KnowledgeRevanth Gangi Reddy, Sai Chetan Chinthakindi, Zhenhailong Wang, Yi R. Fung et al.EMNLP 2022 · 13 citations
- Boosting Data Utilization for Multilingual Dense RetrievalChao Huang, Fengran Mo, Yufeng Chen, Changhao Guan et al.EMNLP 2025 · 2 citations
- Article Reranking by Memory-Enhanced Key Sentence Matching for Detecting Previously Fact-Checked ClaimsQiang Sheng, Juan Cao, Xueyao Zhang, Xirong Li et al.ACL 2021
