Decoding the Rule Book: Extracting Hidden Moderation Criteria from Reddit Communities
Youngwoo Kim, Himanshu Beniwal, Steven L. Johnson, Thomas Hartvigsen
Abstract
Effective content moderation systems require explicit classification criteria, yet online communities like subreddits often operate with diverse, implicit standards. This work introduces a novel approach to identify and extract these implicit criteria from historical moderation data using an interpretable architecture. We represent moderation criteria as score tables of lexical expressions associated with content removal, enabling systematic comparison across different communities. Our experiments demonstrate that these extracted lexical patterns effectively replicate the performance of neural moderation models while providing transparent insights into decision-making processes. The resulting criteria matrix reveals significant variations in how seemingly shared norms are actually enforced, uncovering previously undocumented moderation patterns including community-specific tolerances for language, features for topical restrictions, and underlying subcategories of the toxic speech classification. latex 1 Deleted
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0dd02416-ecdf-47f0-9608-a552595f6decBuilds on4
- "Positive reinforcement helps breed positive behavior": Moderator Perspectives on Encouraging Desirable BehaviorCharlotte Lambert, Frederick Choi, Eshwar ChandrasekharanCSCW 2024 · 9 citations
- Discovering Biases in Information Retrieval Models Using Relevance Thesaurus as Global ExplanationYoungwoo Kim, Razieh Rahimi, James AllanEMNLP 2024 · 3 citations
- ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech DetectionThomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap et al.ACL 2022
- HateCheck: Functional Tests for Hate Speech Detection ModelsPaul Röttger, Bertie Vidgen, Dong Nguyen, Zeerak Waseem et al.ACL 2021
Related papers
- MoMoE: Mixture of Moderation Experts Framework for AI-Assisted Online GovernanceAgam Goyal, Xianyang Zhan, Yilun Chen, Koustuv Saha et al.EMNLP 2025
- Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators through a User-Centric MethodYang Trista Cao, Lovely-Frances Domingo, Sarah A. Gilbert, Michelle L. Mazurek et al.EMNLP 2024 · 4 citations
- RECAST: Enabling User Recourse and Interpretability of Toxicity Detection Models with Interactive VisualizationAustin P. Wright, Omar Shaikh, Haekyu Park, Will Epperson et al.CSCW 2021 · 32 citations
- New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMsShiyao Cui, Qinglin Zhang, Di Wang, Yida Lu et al.ACL 2026
- PluRule: A Benchmark for Moderating Pluralistic Communities on Social MediaZoher Kachwala, Bao Tran Truong, Rasika Muralidharan, Haewoon Kwak et al.ACL 2026
