Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages
Gerrit Quaremba, Amy Rechkemmer, Elizabeth Black, Denny Vrandecic, Elena Simperl
摘要
In automated fact-checking (AFC), check-worthiness detection identifies claims requiring verification based on domain-specific criteria. On Wikipedia, this task instantiates as Citation Needed Detection (CND), which flags claims lacking supporting citations. However, existing research has largely overlooked lower-resource languages, and recent AFC pipelines rely on large language models (LLMs), which are inaccessible to low-resource organizations. We introduce MCN, a multilingual CND corpus spanning 18 languages across three resource levels, on which we conduct an extensive study of small decoder-based language models (SLMs). Our experiments show that SLMs fine-tuned with an encoder-style objective substantially outperform prompted LLMs across languages. We further present one of the first studies on cross-lingual CND, demonstrating that SLMs fine-tuned solely on English claims surpass LLMs, even with little to no target-language adaptation. Our findings have important implications for lower-resource Wikipedia communities and suggest that compact, task-specific models are preferable to LLMs for CND. We release all data and code at https://github.com/gerritq/mcn
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Translation Artifacts in Cross-lingual Transfer LearningMikel Artetxe, Gorka Labaka, Eneko AgirreEMNLP 2020 · 被引用 68 次
- Show Me the Work: Fact-Checkers' Requirements for Explainable Automated Fact-CheckingGreta Warren, Irina Shklovski, Isabelle AugensteinCHI 2025 · 被引用 15 次
- Longitudinal Assessment of Reference Quality on WikipediaAitolkyn Baigutanova, Jaehyeon Myung, Diego Sáez-Trumper, Ai-Jou Chou 等WWW 2023 · 被引用 14 次
相关 Paper
- Lost in Translation, Found in Spans: Identifying Claims in Multilingual Social MediaShubham Mittal, Megha Sundriyal, Preslav NakovEMNLP 2023 · 被引用 4 次
- HintsOfTruth: A Multimodal Checkworthiness Detection Dataset with Real and Synthetic ClaimsMichiel van der Meer, Pavel Korshunov, Sébastien Marcel, Lonneke van der PlasACL 2025 · 被引用 5 次
- An Analysis of Multilingual FActScoreVu Trong Kim, Michael Krumdick, Varshini Reddy, Franck Dernoncourt 等EMNLP 2024 · 被引用 2 次
- X-FACTR: Multilingual Factual Knowledge Retrieval from Pretrained Language ModelsZhengbao Jiang, Antonios Anastasopoulos, Jun Araki, Haibo Ding 等EMNLP 2020 · 被引用 81 次
- ViClaim: A Multilingual Multilabel Dataset for Automatic Claim Detection in VideosPatrick Giedemann, Pius von Däniken, Jan Milan Deriu, Álvaro Rodrigo 等EMNLP 2025 · 被引用 1 次
