HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on Twitter
Manuel Tonneau, Diyi Liu, Niyati Malhotra, Scott A. Hale, Samuel Fraiberger, Víctor Orozco-Olvera, Paul Röttger
Abstract
To address the global challenge of online hate speech, prior research has developed detection models to flag such content on social media. However, due to systematic biases in evaluation datasets, the real-world effectiveness of these models remains unclear, particularly across geographies. We introduce HateDay, the first global hate speech dataset representative of social media settings, constructed from a random sample of all tweets posted on September 21, 2022 and covering eight languages and four English-speaking countries. Using HateDay, we uncover substantial variation in the prevalence and composition of hate speech across languages and regions. We show that evaluations on academic datasets greatly overestimate real-world detection performance, which we find is very low, especially for non-European languages. Our analysis identifies key drivers of this gap, including models' difficulty to distinguish hate from offensive speech and a mismatch between the target groups emphasized in academic datasets and those most frequently targeted in real-world settings. We argue that poor model performance makes public models ill-suited for automatic hate speech moderation and find that high moderation rates are only achievable with substantial human oversight. Our results underscore the need to evaluate detection systems on data that reflects the complexity and diversity of real-world social media.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f79cfc42-ffe5-4da0-801a-65bfc6042ec3Cited by top-tier papers3
- Promptimizer: User-Led Prompt Optimization for Personal Content ClassificationLeijie Wang, Kathryn Yurechko, Amy X. ZhangCHI 2026 · 1 citation
- Compositional Generalisation for Explainable Hate Speech DetectionAgostina Calabrese, Tom Sherborne, Björn Ross, Mirella LapataEMNLP 2025
- TAMA: Target-Aware Multilingual Abuse Detection by Cascaded Conditional Multi-Task LearningJiyan Liu, Youzheng Liu, Taihang Wang, Yimin Wang et al.ACL 2026
Builds on8
- Lost in Moderation: How Commercial Content Moderation APIs Over- and Under-Moderate Group-Targeted Hate Speech and Linguistic VariationsDavid Hartmann, Amin Oueslati, Dimitri Staufer, Lena Pohlmann et al.CHI 2025 · 37 citations
- Measuring the Prevalence of Anti-Social Behavior in Online CommunitiesJoon Sung Park, Joseph Seering, Michael S. BernsteinCSCW 2022 · 24 citations
- On the Challenges of Using Black-Box APIs for Toxicity Evaluation in ResearchLuiza Pozzobon, Beyza Ermis, Patrick Lewis, Sara HookerEMNLP 2023 · 19 citations
- Improving the Detection of Multilingual Online Attacks with Rich Social Media Data from SingaporeJanosch Haber, Bertie Vidgen, Matthew Chapman, Vibhor Agarwal et al.ACL 2023 · 4 citations
- Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language ModelsPaul Röttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck et al.ACL 2024
Related papers
- NaijaHate: Evaluating Hate Speech Detection on Nigerian Twitter Using Representative DataManuel Tonneau, Pedro Vitor Quinta de Castro, Karim Lasri, Ibrahim Farouq et al.ACL 2024
- Spanning the Spectrum of Hatred Detection: A Persian Multi-Label Hate Speech Dataset with Annotator RationalesZahra Delbari, Nafise Sadat Moosavi, Mohammad Taher PilehvarAAAI 2024 · 11 citations
- Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced LanguagesPaul Röttger, Debora Nozza, Federico Bianchi, Dirk HovyEMNLP 2022 · 16 citations
- HateCheck: Functional Tests for Hate Speech Detection ModelsPaul Röttger, Bertie Vidgen, Dong Nguyen, Zeerak Waseem et al.ACL 2021
- Latent Hatred: A Benchmark for Understanding Implicit Hate SpeechMai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi et al.EMNLP 2021 · 159 citations
