An Open Multilingual System for Scoring Readability of Wikipedia
Mykola Trokhymovych, Indira Sen, Martin Gerlach
Abstract
With over 60M articles, Wikipedia has become the largest platform for open and freely accessible knowledge. While it has more than 15B monthly visits, its content is believed to be inaccessible to many readers due to the lack of readability of its text. However, previous investigations of the readability of Wikipedia have been restricted to English only, and there are currently no systems supporting the automatic readability assessment of the 300+ languages in Wikipedia. To bridge this gap, we develop a multilingual model to score the readability of Wikipedia articles. To train and evaluate this model, we create a novel multilingual dataset spanning 14 languages, by matching articles from Wikipedia to simplified Wikipedia and online children encyclopedias. We show that our model performs well in a zero-shot scenario, yielding a ranking accuracy of more than 80% across 14 languages and improving upon previous benchmarks. These results demonstrate the applicability of the model at scale for languages in which there is no ground-truth data available for model fine-tuning. Furthermore, we provide the first overview on the state of readability in Wikipedia beyond English.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 812c962c-8639-40c6-9a04-a5f7409c2201Builds on3
- Neural CRF Model for Sentence Alignment in Text SimplificationChao Jiang, Mounica Maddela, Wuwei Lan, Yang Zhong et al.ACL 2020 · 103 citations
- Document-Level Text Simplification: Dataset, Criteria and BaselineRenliang Sun, Hanqi Jin, Xiaojun WanEMNLP 2021 · 33 citations
- SWiPE: A Dataset for Document-Level Simplification of Wikipedia PagesPhilippe Laban, Jesse Vig, Wojciech Kryscinski, Shafiq Joty et al.ACL 2023 · 5 citations
Related papers
- Crosslingual Topic Modeling with WikiPDATiziano Piccardi, Robert WestWWW 2021 · 18 citations
- How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLPKushal Tatariya, Artur Kulmizev, Wessel Poelman, Esther Ploeger et al.ACL 2026 · 2 citations
- XWikiGen: Cross-lingual Summarization for Encyclopedic Text Generation in Low Resource LanguagesDhaval Taunk, Shivprasad Sagare, Anupam Patil, Shivansh Subramanian et al.WWW 2023 · 3 citations
- Descartes: Generating Short Descriptions of Wikipedia ArticlesMarija Sakota, Maxime Peyrard, Robert WestWWW 2023 · 6 citations
- Models and Datasets for Cross-Lingual SummarisationLaura Perez-Beltrachini, Mirella LapataEMNLP 2021 · 1 citation
