WikiSQE: A Large-Scale Dataset for Sentence Quality Estimation in Wikipedia
Kenichiro Ando, Satoshi Sekine, Mamoru Komachi
Abstract
Wikipedia can be edited by anyone and thus contains various quality sentences. Therefore, Wikipedia includes some poor-quality edits, which are often marked up by other editors. While editors' reviews enhance the credibility of Wikipedia, it is hard to check all edited text. Assisting in this process is very important, but a large and comprehensive dataset for studying it does not currently exist. Here, we propose WikiSQE, the first large-scale dataset for sentence quality estimation in Wikipedia. Each sentence is extracted from the entire revision history of English Wikipedia, and the target quality labels were carefully investigated and selected. WikiSQE has about 3.4 M sentences with 153 quality labels. In the experiment with automatic classification using competitive machine learning models, sentences that had problems with citation, syntax/semantics, or propositions were found to be more difficult to detect. In addition, by performing human annotation, we found that the model we developed performed better than the crowdsourced workers. WikiSQE is expected to be a valuable resource for other tasks in NLP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3c3a96b-3598-45b4-b61f-e5c3ed5398e7Builds on4
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
- ParaCrawl: Web-Scale Acquisition of Parallel CorporaMarta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield et al.ACL 2020 · 132 citations
- Automatically Labeling Low Quality Content on Wikipedia By Leveraging Patterns in Editing BehaviorsSumit Asthana, Sabrina Tobar Thommel, Aaron Lee Halfaker, Nikola BanovicCSCW 2021 · 9 citations
- StereoSet: Measuring stereotypical bias in pretrained language modelsMoin Nadeem, Anna Bethke, Siva ReddyACL 2021
Related papers
- How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLPKushal Tatariya, Artur Kulmizev, Wessel Poelman, Esther Ploeger et al.ACL 2026 · 2 citations
- Quality Change: Norm or Exception? Measurement, Analysis and Detection of Quality Change in WikipediaParamita Das, Bhanu Prakash Reddy Guda, Sasi Bhushan Seelaboyina, Soumya Sarkar et al.CSCW 2022 · 8 citations
- NwQM: A neural quality assessment framework for WikipediaBhanu Prakash Reddy Guda, Sasi Bhushan Seelaboyina, Soumya Sarkar, Animesh MukherjeeEMNLP 2020 · 7 citations
- Re-TACRED: Addressing Shortcomings of the TACRED DatasetGeorge Stoica, Emmanouil Antonios Platanios, Barnabás PóczosAAAI 2021 · 146 citations
- Longitudinal Assessment of Reference Quality on WikipediaAitolkyn Baigutanova, Jaehyeon Myung, Diego Sáez-Trumper, Ai-Jou Chou et al.WWW 2023 · 14 citations
