Why Should This Article Be Deleted? Transparent Stance Detection in Multilingual Wikipedia Editor Discussions
Lucie-Aimée Kaffee, Arnav Arora, Isabelle Augenstein
Abstract
The moderation of content on online platforms is usually non-transparent. On Wikipedia, however, this discussion is carried out publicly and editors are encouraged to use the content moderation policies as explanations for making moderation decisions. Currently, only a few comments explicitly mention those policies -20% of the English ones, but as few as 2% of the German and Turkish comments. To aid in this process of understanding how content is moderated, we construct a novel multilingual dataset of Wikipedia editor discussions along with their reasoning in three languages. The dataset contains the stances of the editors (keep, delete, merge, comment), along with the stated reason, and a content moderation policy, for each edit decision. We demonstrate that stance and corresponding reason (policy) can be predicted jointly with a high degree of accuracy, adding transparency to the decision-making process. We release both our joint prediction models and the multilingual content moderation dataset for further research on automated transparent content moderation. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5a7599ad-b545-42fc-be2c-8d87b49c41c3Builds on7
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Generating Fact Checking ExplanationsPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinACL 2020 · 130 citations
- Few-Shot Cross-Lingual Stance Detection with Sentiment-Based Pre-trainingMomchil Hardalov, Arnav Arora, Preslav Nakov, Isabelle AugensteinAAAI 2022 · 72 citations
- Automated Content Moderation Increases Adherence to Community GuidelinesManoel Horta Ribeiro, Justin Cheng, Robert WestWWW 2023 · 53 citations
- The State and Fate of Linguistic Diversity and Inclusion in the NLP WorldPratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali et al.ACL 2020 · 40 citations
Related papers
- PerspectiveMod: A Perspectivist Resource for Deliberative ModerationEva Maria Vecchi, Neele Falk, Carlotta Quensel, Iman Jundi et al.EMNLP 2025
- Wikipedia Beyond the English Language Edition: How do Editors Collaborate in the Farsi and Chinese Wikipedias?Taryn Bipat, Negin Alimohammadi, Yihan Yu, David W. McDonald et al.CSCW 2021 · 3 citations
- ICM-Assistant: Instruction-tuning Multimodal Large Language Models for Rule-based Explainable Image Content ModerationMengyang Wu, Yuzhi Zhao, Jialun Cao, Mingjie Xu et al.AAAI 2025 · 14 citations
- Wikibench: Community-Driven Data Curation for AI Evaluation on WikipediaTzu-Sheng Kuo, Aaron Lee Halfaker, Zirui Cheng, Jiwoo Kim et al.CHI 2024 · 19 citations
- An Open Multilingual System for Scoring Readability of WikipediaMykola Trokhymovych, Indira Sen, Martin GerlachACL 2024
