Decoding the Rule Book: Extracting Hidden Moderation Criteria from Reddit Communities
Youngwoo Kim, Himanshu Beniwal, Steven L. Johnson, Thomas Hartvigsen
摘要
Effective content moderation systems require explicit classification criteria, yet online communities like subreddits often operate with diverse, implicit standards. This work introduces a novel approach to identify and extract these implicit criteria from historical moderation data using an interpretable architecture. We represent moderation criteria as score tables of lexical expressions associated with content removal, enabling systematic comparison across different communities. Our experiments demonstrate that these extracted lexical patterns effectively replicate the performance of neural moderation models while providing transparent insights into decision-making processes. The resulting criteria matrix reveals significant variations in how seemingly shared norms are actually enforced, uncovering previously undocumented moderation patterns including community-specific tolerances for language, features for topical restrictions, and underlying subcategories of the toxic speech classification. latex 1 Deleted
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- "Positive reinforcement helps breed positive behavior": Moderator Perspectives on Encouraging Desirable BehaviorCharlotte Lambert, Frederick Choi, Eshwar ChandrasekharanCSCW 2024 · 被引用 9 次
- Discovering Biases in Information Retrieval Models Using Relevance Thesaurus as Global ExplanationYoungwoo Kim, Razieh Rahimi, James AllanEMNLP 2024 · 被引用 3 次
- ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech DetectionThomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap 等ACL 2022
- HateCheck: Functional Tests for Hate Speech Detection ModelsPaul Röttger, Bertie Vidgen, Dong Nguyen, Zeerak Waseem 等ACL 2021
相关 Paper
- MoMoE: Mixture of Moderation Experts Framework for AI-Assisted Online GovernanceAgam Goyal, Xianyang Zhan, Yilun Chen, Koustuv Saha 等EMNLP 2025
- Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators through a User-Centric MethodYang Trista Cao, Lovely-Frances Domingo, Sarah A. Gilbert, Michelle L. Mazurek 等EMNLP 2024 · 被引用 4 次
- RECAST: Enabling User Recourse and Interpretability of Toxicity Detection Models with Interactive VisualizationAustin P. Wright, Omar Shaikh, Haekyu Park, Will Epperson 等CSCW 2021 · 被引用 32 次
- New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMsShiyao Cui, Qinglin Zhang, Di Wang, Yida Lu 等ACL 2026
- PluRule: A Benchmark for Moderating Pluralistic Communities on Social MediaZoher Kachwala, Bao Tran Truong, Rasika Muralidharan, Haewoon Kwak 等ACL 2026
