An Open Multilingual System for Scoring Readability of Wikipedia
Mykola Trokhymovych, Indira Sen, Martin Gerlach
摘要
With over 60M articles, Wikipedia has become the largest platform for open and freely accessible knowledge. While it has more than 15B monthly visits, its content is believed to be inaccessible to many readers due to the lack of readability of its text. However, previous investigations of the readability of Wikipedia have been restricted to English only, and there are currently no systems supporting the automatic readability assessment of the 300+ languages in Wikipedia. To bridge this gap, we develop a multilingual model to score the readability of Wikipedia articles. To train and evaluate this model, we create a novel multilingual dataset spanning 14 languages, by matching articles from Wikipedia to simplified Wikipedia and online children encyclopedias. We show that our model performs well in a zero-shot scenario, yielding a ranking accuracy of more than 80% across 14 languages and improving upon previous benchmarks. These results demonstrate the applicability of the model at scale for languages in which there is no ground-truth data available for model fine-tuning. Furthermore, we provide the first overview on the state of readability in Wikipedia beyond English.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Neural CRF Model for Sentence Alignment in Text SimplificationChao Jiang, Mounica Maddela, Wuwei Lan, Yang Zhong 等ACL 2020 · 被引用 103 次
- Document-Level Text Simplification: Dataset, Criteria and BaselineRenliang Sun, Hanqi Jin, Xiaojun WanEMNLP 2021 · 被引用 33 次
- SWiPE: A Dataset for Document-Level Simplification of Wikipedia PagesPhilippe Laban, Jesse Vig, Wojciech Kryscinski, Shafiq Joty 等ACL 2023 · 被引用 5 次
相关 Paper
- Crosslingual Topic Modeling with WikiPDATiziano Piccardi, Robert WestWWW 2021 · 被引用 18 次
- How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLPKushal Tatariya, Artur Kulmizev, Wessel Poelman, Esther Ploeger 等ACL 2026 · 被引用 2 次
- XWikiGen: Cross-lingual Summarization for Encyclopedic Text Generation in Low Resource LanguagesDhaval Taunk, Shivprasad Sagare, Anupam Patil, Shivansh Subramanian 等WWW 2023 · 被引用 3 次
- Descartes: Generating Short Descriptions of Wikipedia ArticlesMarija Sakota, Maxime Peyrard, Robert WestWWW 2023 · 被引用 6 次
- Models and Datasets for Cross-Lingual SummarisationLaura Perez-Beltrachini, Mirella LapataEMNLP 2021 · 被引用 1 次
