How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP
Kushal Tatariya, Artur Kulmizev, Wessel Poelman, Esther Ploeger, Marcel Bollmann, Johannes Bjerva, Jiaming Luo, Heather C. Lent, Miryam de Lhoneux
Abstract
Wikipedia's perceived high quality and broad language coverage have established it as a fundamental resource in NLP. However, in recent years, such assumptions of high quality have become the subject of scrutiny in low-resource and multilingual contexts. In this study, we subject the entirety of non-English Wikipedia to a data filtering procedure typically reserved for noisy web-text -- a process which removes a large percentage of the collection's data. In analysing the removed data, we reveal numerous systematic quality issues, such as script and language contamination, repeated template and placeholder articles, and a high concentration of bot-generated content. We consolidate these findings into a 4-level quality ranking of Wikipedia, which shows strong correspondence with alternative quality measures and heuristics. Lastly, we evaluate the downstream impact of quality filtering in three practical language modelling scenarios, showing that models trained on filtered data largely match or outperform those trained on raw Wikipedia, with the largest gains observed for lower-quality language editions. Ultimately, our experiments serve as a first step in establishing quality-aware best practices for Wikipedia utilization in NLP, laying groundwork that can inform future dataset creation and curation efforts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 31427dcb-9fd2-4d48-9280-fe1d98bcce17Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- On the Cross-lingual Transferability of Monolingual RepresentationsMikel Artetxe, Sebastian Ruder, Dani YogatamaACL 2020 · 57 citations
- MLQA: Evaluating Cross-lingual Extractive Question AnsweringPatrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel et al.ACL 2020 · 52 citations
- MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual TransferJonas Pfeiffer, Ivan Vulic, Iryna Gurevych, Sebastian RuderEMNLP 2020 · 36 citations
Related papers
- Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data SelectionSuchin Gururangan, Dallas Card, Sarah K. Dreier, Emily K. Gade et al.EMNLP 2022 · 41 citations
- WikiSQE: A Large-Scale Dataset for Sentence Quality Estimation in WikipediaKenichiro Ando, Satoshi Sekine, Mamoru KomachiAAAI 2024 · 5 citations
- NwQM: A neural quality assessment framework for WikipediaBhanu Prakash Reddy Guda, Sasi Bhushan Seelaboyina, Soumya Sarkar, Animesh MukherjeeEMNLP 2020 · 7 citations
- An Open Multilingual System for Scoring Readability of WikipediaMykola Trokhymovych, Indira Sen, Martin GerlachACL 2024
- An Analysis of Multilingual FActScoreVu Trong Kim, Michael Krumdick, Varshini Reddy, Franck Dernoncourt et al.EMNLP 2024 · 2 citations
