Lost in Translation, Found in Spans: Identifying Claims in Multilingual Social Media
Shubham Mittal, Megha Sundriyal, Preslav Nakov
摘要
Claim span identification (CSI) is an important step in fact-checking pipelines, aiming to identify text segments that contain a check-worthy claim or assertion in a social media post. Despite its importance to journalists and human fact-checkers, it remains a severely understudied problem, and the scarce research on this topic so far has only focused on English. Here we aim to bridge this gap by creating a novel dataset, X-CLAIM, consisting of 7K real-world claims collected from numerous social media platforms in five Indian languages and English. We report strong baselines with state-of-the-art encoder-only language models (e.g., XLM-R) and we demonstrate the benefits of training on multiple languages over alternative cross-lingual transfer methods such as zero-shot transfer, or training on translated data, from a high-resource language such as English. We evaluate generative large language models from the GPT series using prompting methods on the X-CLAIM dataset and we find that they underperform the smaller encoder-only language models for low-resource languages.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 被引用 394 次
- Alignment-Augmented Consistent Translation for Multilingual Open Information ExtractionKeshav Kolluru, Muqeeth Mohammed, Shubham Mittal, Soumen Chakrabarti 等ACL 2022
- Everything Is All It Takes: A Multipronged Strategy for Zero-Shot Cross-Lingual Information ExtractionMahsa Yarmohammadi, Shijie Wu, Marc Marone, Haoran Xu 等EMNLP 2021
相关 Paper
- Multilingual Previously Fact-Checked Claim RetrievalMatús Pikuliak, Ivan Srba, Róbert Móro, Timo Hromadka 等EMNLP 2023 · 被引用 9 次
- ViClaim: A Multilingual Multilabel Dataset for Automatic Claim Detection in VideosPatrick Giedemann, Pius von Däniken, Jan Milan Deriu, Álvaro Rodrigo 等EMNLP 2025 · 被引用 1 次
- Claim Matching Beyond English to Scale Global Fact-CheckingAshkan Kazemi, Kiran Garimella, Devin Gaffney, Scott HaleACL 2021
- AFaCTA: Assisting the Annotation of Factual Claim Detection with Reliable LLM AnnotatorsJingwei Ni, Minjing Shi, Dominik Stammbach, Mrinmaya Sachan 等ACL 2024
- Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource LanguagesGerrit Quaremba, Amy Rechkemmer, Elizabeth Black, Denny Vrandecic 等ACL 2026
