Lost in Translation, Found in Spans: Identifying Claims in Multilingual Social Media
Shubham Mittal, Megha Sundriyal, Preslav Nakov
Abstract
Claim span identification (CSI) is an important step in fact-checking pipelines, aiming to identify text segments that contain a check-worthy claim or assertion in a social media post. Despite its importance to journalists and human fact-checkers, it remains a severely understudied problem, and the scarce research on this topic so far has only focused on English. Here we aim to bridge this gap by creating a novel dataset, X-CLAIM, consisting of 7K real-world claims collected from numerous social media platforms in five Indian languages and English. We report strong baselines with state-of-the-art encoder-only language models (e.g., XLM-R) and we demonstrate the benefits of training on multiple languages over alternative cross-lingual transfer methods such as zero-shot transfer, or training on translated data, from a high-resource language such as English. We evaluate generative large language models from the GPT series using prompting methods on the X-CLAIM dataset and we find that they underperform the smaller encoder-only language models for low-resource languages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 523de1da-3a5e-4a00-bc9b-aa24c8465521Cited by top-tier papers1
Ask how each one uses itBuilds on5
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
- Alignment-Augmented Consistent Translation for Multilingual Open Information ExtractionKeshav Kolluru, Muqeeth Mohammed, Shubham Mittal, Soumen Chakrabarti et al.ACL 2022
- Everything Is All It Takes: A Multipronged Strategy for Zero-Shot Cross-Lingual Information ExtractionMahsa Yarmohammadi, Shijie Wu, Marc Marone, Haoran Xu et al.EMNLP 2021
Related papers
- Multilingual Previously Fact-Checked Claim RetrievalMatús Pikuliak, Ivan Srba, Róbert Móro, Timo Hromadka et al.EMNLP 2023 · 9 citations
- ViClaim: A Multilingual Multilabel Dataset for Automatic Claim Detection in VideosPatrick Giedemann, Pius von Däniken, Jan Milan Deriu, Álvaro Rodrigo et al.EMNLP 2025 · 1 citation
- Claim Matching Beyond English to Scale Global Fact-CheckingAshkan Kazemi, Kiran Garimella, Devin Gaffney, Scott HaleACL 2021
- AFaCTA: Assisting the Annotation of Factual Claim Detection with Reliable LLM AnnotatorsJingwei Ni, Minjing Shi, Dominik Stammbach, Mrinmaya Sachan et al.ACL 2024
- Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource LanguagesGerrit Quaremba, Amy Rechkemmer, Elizabeth Black, Denny Vrandecic et al.ACL 2026
