NaijaHate: Evaluating Hate Speech Detection on Nigerian Twitter Using Representative Data
Manuel Tonneau, Pedro Vitor Quinta de Castro, Karim Lasri, Ibrahim Farouq, Lakshmi Subramanian, Víctor Orozco-Olvera, Samuel Fraiberger
Abstract
To address the global issue of online hate, hate speech detection (HSD) systems are typically developed on datasets from the United States, thereby failing to generalize to English dialects from the Majority World. Furthermore, HSD models are often evaluated on non-representative samples, raising concerns about overestimating model performance in real-world settings. In this work, we introduce NAIJAHATE, the first dataset annotated for HSD which contains a representative sample of Nigerian tweets. We demonstrate that HSD evaluated on biased datasets traditionally used in the literature consistently overestimates real-world performance by at least two-fold. We then propose NAIJAXLM-T, a pretrained model tailored to the Nigerian Twitter context, and establish the key role played by domainadaptive pretraining and finetuning in maximizing HSD performance. Finally, owing to the modest performance of HSD systems in realworld conditions, we find that content moderators would need to review about ten thousand Nigerian tweets flagged as hateful daily to moderate 60% of all hateful content, highlighting the challenges of moderating hate speech at scale as social media usage continues to grow globally. Taken together, these results pave the way towards robust HSD systems and a better protection of social media users from hateful content in low-resource settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb014fda-5b2c-4c48-8b8d-0ac8984b8e70Cited by top-tier papers4
- HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on TwitterManuel Tonneau, Diyi Liu, Niyati Malhotra, Scott A. Hale et al.ACL 2025 · 12 citations
- On the Risk of Evidence Pollution for Malicious Social Text Detection in the Era of LLMsHerun Wan, Minnan Luo, Zhixiong Su, Guang Dai et al.ACL 2025 · 5 citations
- SAGE: Synergistic Adaptive Gating of Experts for Hateful Video DetectionJie Huang, Xin Liao, Junjie Wang, Mingyang Li et al.ACL 2026
- Culture Cartography: Mapping the Landscape of Cultural KnowledgeCaleb Ziems, William Barr Held, Jane Yu, Amir Goldberg et al.EMNLP 2025
Builds on8
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
- The Psychological Well-Being of Content Moderators: The Emotional Labor of Commercial Moderation and Avenues for Improving SupportMiriah Steiger, Timir J. Bharucha, Sukrit Venkatagiri, Martin J. Riedl et al.CHI 2021 · 168 citations
- Human-AI Collaboration via Conditional Delegation: A Case Study of Content ModerationVivian Lai, Samuel Carton, Rajat Bhatnagar, Q. Vera Liao et al.CHI 2022 · 135 citations
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo et al.ACL 2020 · 93 citations
Related papers
- Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced LanguagesPaul Röttger, Debora Nozza, Federico Bianchi, Dirk HovyEMNLP 2022 · 16 citations
- LLM-Based Multi-Task Bangla Hate Speech Detection: Type, Severity, and TargetMd. Arid Hasan, Firoj Alam, Md Fahad Hossain, Usman Naseem et al.ACL 2026
- SAHSD: Enhancing Hate Speech Detection in LLM-Powered Web Applications via Sentiment Analysis and Few-Shot LearningYulong Wang, Hong Li, Ni WeiWWW 2025 · 2 citations
- Spanning the Spectrum of Hatred Detection: A Persian Multi-Label Hate Speech Dataset with Annotator RationalesZahra Delbari, Nafise Sadat Moosavi, Mohammad Taher PilehvarAAAI 2024 · 11 citations
- Pinpointing Fine-Grained Relationships between Hateful Tweets and RepliesAbdullah Albanyan, Eduardo BlancoAAAI 2022 · 9 citations
