Rule By Example: Harnessing Logical Rules for Explainable Hate Speech Detection
Christopher Clarke, Matthew Hall, Gaurav Mittal, Ye Yu, Sandra Sajeev, Jason Mars, Mei Chen
Abstract
Classic approaches to content moderation typically apply a rule-based heuristic approach to flag content. While rules are easily customizable and intuitive for humans to interpret, they are inherently fragile and lack the flexibility or robustness needed to moderate the vast amount of undesirable content found online today. Recent advances in deep learning have demonstrated the promise of using highly effective deep neural models to overcome these challenges. However, despite the improved performance, these data-driven models lack transparency and explainability, often leading to mistrust from everyday users and a lack of adoption by many platforms. In this paper, we present Rule By Example (RBE): a novel exemplarbased contrastive learning approach for learning from logical rules for the task of textual content moderation. RBE is capable of providing rule-grounded predictions, allowing for more explainable and customizable predictions compared to typical deep learning-based approaches. We demonstrate that our approach is capable of learning rich rule embedding representations using only a few data examples. Experimental results on 3 popular hate speech classification datasets show that RBE is able to outperform state-of-the-art deep learning classifiers as well as the use of rules in both supervised and unsupervised settings while providing explainable model predictions via rulegrounding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ffb1c1aa-d5b4-4e59-931f-af9889a8769cCited by top-tier papers2
- Efficient LLM Moderation with Multi-Layer Latent PrototypesMaciej Chrabaszcz, Filip Szatkowski, Bartosz Wójcik, Jan Dubiński et al.ICML 2026
- Machines in the Margins: A Systematic Review of Automated Content Generation for WikipediaNeal Reeves, Elena SimperlCSCW 2025
Builds on6
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- Disproportionate Removals and Differing Content Moderation Experiences for Conservative, Transgender, and Black Social Media Users: Marginalization and Moderation Gray AreasOliver L. Haimson, Daniel Delmonaco, Peipei Nie, Andrea WegnerCSCW 2021 · 287 citations
- Human-AI Collaboration via Conditional Delegation: A Case Study of Content ModerationVivian Lai, Samuel Carton, Rajat Bhatnagar, Q. Vera Liao et al.CHI 2022 · 135 citations
- Learning from Rules Generalizing Labeled ExemplarsAbhijeet Awasthi, Sabyasachi Ghosh, Rasna Goyal, Sunita SarawagiICLR 2020 · 93 citations
Related papers
- CHAIRO: Contextual Hierarchical Analogical Induction and Reasoning Optimization for LLMsHaotian Lu, Yuchen Mou, Bingzhe WuACL 2026
- Interpretable Neural-Symbolic Concept ReasoningPietro Barbiero, Gabriele Ciravegna, Francesco Giannini, Mateo Espinosa Zarlenga et al.ICML 2023 · 68 citations
- Improving Hateful Meme Detection through Retrieval-Guided Contrastive LearningJingbiao Mei, Jinghong Chen, Weizhe Lin, Bill Byrne et al.ACL 2024 · 13 citations
- MoMoE: Mixture of Moderation Experts Framework for AI-Assisted Online GovernanceAgam Goyal, Xianyang Zhan, Yilun Chen, Koustuv Saha et al.EMNLP 2025
- AmpleHate: Amplifying the Attention for Versatile Implicit Hate DetectionYejin Lee, Joonghyuk Hahn, Hyeseon Ahn, Yo-Sub HanEMNLP 2025 · 2 citations
