Rule By Example: Harnessing Logical Rules for Explainable Hate Speech Detection
Christopher Clarke, Matthew Hall, Gaurav Mittal, Ye Yu, Sandra Sajeev, Jason Mars, Mei Chen
摘要
Classic approaches to content moderation typically apply a rule-based heuristic approach to flag content. While rules are easily customizable and intuitive for humans to interpret, they are inherently fragile and lack the flexibility or robustness needed to moderate the vast amount of undesirable content found online today. Recent advances in deep learning have demonstrated the promise of using highly effective deep neural models to overcome these challenges. However, despite the improved performance, these data-driven models lack transparency and explainability, often leading to mistrust from everyday users and a lack of adoption by many platforms. In this paper, we present Rule By Example (RBE): a novel exemplarbased contrastive learning approach for learning from logical rules for the task of textual content moderation. RBE is capable of providing rule-grounded predictions, allowing for more explainable and customizable predictions compared to typical deep learning-based approaches. We demonstrate that our approach is capable of learning rich rule embedding representations using only a few data examples. Experimental results on 3 popular hate speech classification datasets show that RBE is able to outperform state-of-the-art deep learning classifiers as well as the use of rules in both supervised and unsupervised settings while providing explainable model predictions via rulegrounding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Efficient LLM Moderation with Multi-Layer Latent PrototypesMaciej Chrabaszcz, Filip Szatkowski, Bartosz Wójcik, Jan Dubiński 等ICML 2026
- Machines in the Margins: A Systematic Review of Automated Content Generation for WikipediaNeal Reeves, Elena SimperlCSCW 2025
它引用的顶会 Paper6
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
- Disproportionate Removals and Differing Content Moderation Experiences for Conservative, Transgender, and Black Social Media Users: Marginalization and Moderation Gray AreasOliver L. Haimson, Daniel Delmonaco, Peipei Nie, Andrea WegnerCSCW 2021 · 被引用 287 次
- Human-AI Collaboration via Conditional Delegation: A Case Study of Content ModerationVivian Lai, Samuel Carton, Rajat Bhatnagar, Q. Vera Liao 等CHI 2022 · 被引用 135 次
- Learning from Rules Generalizing Labeled ExemplarsAbhijeet Awasthi, Sabyasachi Ghosh, Rasna Goyal, Sunita SarawagiICLR 2020 · 被引用 93 次
相关 Paper
- CHAIRO: Contextual Hierarchical Analogical Induction and Reasoning Optimization for LLMsHaotian Lu, Yuchen Mou, Bingzhe WuACL 2026
- Interpretable Neural-Symbolic Concept ReasoningPietro Barbiero, Gabriele Ciravegna, Francesco Giannini, Mateo Espinosa Zarlenga 等ICML 2023 · 被引用 68 次
- Improving Hateful Meme Detection through Retrieval-Guided Contrastive LearningJingbiao Mei, Jinghong Chen, Weizhe Lin, Bill Byrne 等ACL 2024 · 被引用 13 次
- MoMoE: Mixture of Moderation Experts Framework for AI-Assisted Online GovernanceAgam Goyal, Xianyang Zhan, Yilun Chen, Koustuv Saha 等EMNLP 2025
- AmpleHate: Amplifying the Attention for Versatile Implicit Hate DetectionYejin Lee, Joonghyuk Hahn, Hyeseon Ahn, Yo-Sub HanEMNLP 2025 · 被引用 2 次
