Ruddit: Norms of Offensiveness for English Reddit Comments
Rishav Hada, Sohi Sudhir, Pushkar Mishra, Helen Yannakoudakis, Saif M. Mohammad, Ekaterina Shutova
Abstract
Warning: This paper contains comments that may be offensive or upsetting. On social media platforms, hateful and offensive language negatively impact the mental well-being of users and the participation of people from diverse backgrounds. Automatic methods to detect offensive language have largely relied on datasets with categorical labels. However, comments can vary in their degree of offensiveness. We create the first dataset of English language Reddit comments that has finegrained, real-valued scores between -1 (maximally supportive) and 1 (maximally offensive). The dataset was annotated using Best-Worst Scaling, a form of comparative annotation that has been shown to alleviate known biases of using rating scales. We show that the method produces highly reliable offensiveness scores. Finally, we evaluate the ability of widely-used neural models to predict offensiveness scores on this new dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c5459d8a-6f0e-4af5-99ee-8e353c46abcaCited by top-tier papers6
- Just Say No: Analyzing the Stance of Neural Dialogue Generation in Offensive ContextsAshutosh Baheti, Maarten Sap, Alan Ritter, Mark O. RiedlEMNLP 2021 · 50 citations
- Which Demographics do LLMs Default to During Annotation?Johannes Schäfer, Aidan Combs, Christopher Bagdon, Jiahui Li et al.ACL 2025 · 11 citations
- "Fifty Shades of Bias": Normative Ratings of Gender Bias in GPT Generated English TextRishav Hada, Agrima Seth, Harshita Diddee, Kalika BaliEMNLP 2023 · 10 citations
- End User Authoring of Personalized Content Classifiers: Comparing Example Labeling, Rule Writing, and LLM PromptingLeijie Wang, Kathryn Yurechko, Pranati Dani, Quan Ze Chen et al.CHI 2025 · 7 citations
- HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard ModelsSeanie Lee, Haebin Seong, Dong Bok Lee, Minki Kang et al.ICLR 2025
Builds on1
Related papers
- Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators' DisagreementElisa Leonardelli, Stefano Menini, Alessio Palmero Aprosio, Marco Guerini et al.EMNLP 2021 · 2 citations
- KOLD: Korean Offensive Language DatasetYounghoon Jeong, Juhyun Oh, Jongwon Lee, Jaimeen Ahn et al.EMNLP 2022 · 41 citations
- Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is OffensiveTharindu Cyril Weerasooriya, Sujan Dutta, Tharindu Ranasinghe, Marcos Zampieri et al.EMNLP 2023 · 12 citations
- GoEmotions: A Dataset of Fine-Grained EmotionsDorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan S. Cowen et al.ACL 2020 · 16 citations
- Pinpointing Fine-Grained Relationships between Hateful Tweets and RepliesAbdullah Albanyan, Eduardo BlancoAAAI 2022 · 9 citations
