Automatically Labeling Low Quality Content on Wikipedia By Leveraging Patterns in Editing Behaviors
Sumit Asthana, Sabrina Tobar Thommel, Aaron Lee Halfaker, Nikola Banovic
Abstract
Wikipedia articles aim to be definitive sources of encyclopedic content. Yet, only 0.6% of Wikipedia articles have high quality according to its quality scale due to insufficient number of Wikipedia editors and enormous number of articles. Supervised Machine Learning (ML) quality improvement approaches that can automatically identify and fix content issues rely on manual labels of individual Wikipedia sentence quality. However, current labeling approaches are tedious and produce noisy labels. Here, we propose an automated labeling approach that identifies the semantic category (e.g., adding citations, clarifications) of historic Wikipedia edits and uses the modified sentences prior to the edit as examples that require that semantic improvement. Highest-rated article sentences are examples that no longer need semantic improvements. We show that training existing sentence quality classification algorithms on our labels improves their performance compared to training them on existing labels. Our work shows that editing behaviors of Wikipedia editors provide better labels than labels generated by crowdworkers who lack the context to make judgments that the editors would agree with.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d86a087b-c646-4bd9-88f0-d6c0cab69f40Cited by top-tier papers3
- WikiSQE: A Large-Scale Dataset for Sentence Quality Estimation in WikipediaKenichiro Ando, Satoshi Sekine, Mamoru KomachiAAAI 2024 · 5 citations
- To Revise or Not to Revise: Learning to Detect Improvable Claims for Argumentative Writing SupportGabriella Skitalinskaya, Henning WachsmuthACL 2023 · 2 citations
- Machines in the Margins: A Systematic Review of Automated Content Generation for WikipediaNeal Reeves, Elena SimperlCSCW 2025
Related papers
- NwQM: A neural quality assessment framework for WikipediaBhanu Prakash Reddy Guda, Sasi Bhushan Seelaboyina, Soumya Sarkar, Animesh MukherjeeEMNLP 2020 · 7 citations
- Quality Change: Norm or Exception? Measurement, Analysis and Detection of Quality Change in WikipediaParamita Das, Bhanu Prakash Reddy Guda, Sasi Bhushan Seelaboyina, Soumya Sarkar et al.CSCW 2022 · 8 citations
- How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLPKushal Tatariya, Artur Kulmizev, Wessel Poelman, Esther Ploeger et al.ACL 2026 · 2 citations
- Descartes: Generating Short Descriptions of Wikipedia ArticlesMarija Sakota, Maxime Peyrard, Robert WestWWW 2023 · 6 citations
- Longitudinal Assessment of Reference Quality on WikipediaAitolkyn Baigutanova, Jaehyeon Myung, Diego Sáez-Trumper, Ai-Jou Chou et al.WWW 2023 · 14 citations
