When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks
Eve Fleisig, Rediet Abebe, Dan Klein
Abstract
Though majority vote among annotators is typically used for ground truth labels in machine learning, annotator disagreement in tasks such as hate speech detection may reflect systematic differences in opinion across groups, not noise. Thus, a crucial problem in hate speech detection is determining if a statement is offensive to the demographic group that it targets, when that group may be a small fraction of the annotator pool. We construct a model that predicts individual annotator ratings on potentially offensive text and combines this information with the predicted target group of the text to predict the ratings of target group members. We show gains across a range of metrics, including raising performance over the baseline by 22% at predicting individual annotators' ratings and by 33% at predicting variance among annotators, which provides a metric for model uncertainty downstream. We find that annotators' ratings can be predicted using their demographic information as well as opinions on online content, and that non-invasive questions on annotators' online experiences minimize the need to collect demographic information when predicting annotators' opinions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 67d5b29d-660a-4219-b315-139ddd61a5edCited by top-tier papers24
- Personalizing Reinforcement Learning from Human Feedback with Variational Preference LearningSriyash Poddar, Yanming Wan, Hamish Ivison, Abhishek Gupta et al.NeurIPS 2024 · 188 citations
- Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHFAnand Siththaranjan, Cassidy Laidlaw, Dylan Hadfield-MenellICLR 2024 · 112 citations
- WildFeedback: Aligning LLMs With In-situ User Interactions And FeedbackTaiwei Shi, Zhuoer Wang, Longqi Yang, Ying-Chun Lin et al.ACL 2026 · 35 citations
- Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text PerceptionsMatthias Orlikowski, Jiaxin Pei, Paul Röttger, Philipp Cimiano et al.ACL 2025 · 34 citations
- Improving Context-Aware Preference Modeling for Language ModelsSilviu Pitis, Ziang Xiao, Nicolas Le Roux, Alessandro SordoniNeurIPS 2024 · 30 citations
Builds on4
- Jury Learning: Integrating Dissenting Voices into Machine Learning ModelsMitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel et al.CHI 2022 · 134 citations
- Social Bias Frames: Reasoning about Social and Power Implications of LanguageMaarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky et al.ACL 2020 · 16 citations
- FairPrism: Evaluating Fairness-Related Harms in Text GenerationEve Fleisig, Aubrie Amstutz, Chad Atalla, Su Lin Blodgett et al.ACL 2023 · 9 citations
- Unifying Data Perspectivism and Personalization: An Application to Social NormsJoan Plepi, Béla Neuendorf, Lucie Flek, Charles WelchEMNLP 2022 · 4 citations
Related papers
- Impact of Annotator Demographics on Sentiment Dataset LabelingYi Ding, Jacob You, Tonja-Katrin Machulla, Jennifer Jacobs et al.CSCW 2022 · 20 citations
- Disentangling Subjectivity and Uncertainty for Hate Speech Annotation and Modeling using GazeÖzge Alaçam, Sanne Hoeken, Andreas Säuberli, Hannes Gröner et al.EMNLP 2025
- From Granular Grief to Binary Belief: A Collaborative Optimization of Annotation Techniques for Anti-Autistic LanguageNaba Rizvi, Alexis Morales Flores, Mohammad Rizvi, Nedjma Ousidhoum et al.CSCW 2025 · 2 citations
- PREDICT: Multi-Agent-based Debate Simulation for Generalized Hate Speech DetectionSomeen Park, Jaehoon Kim, Seungwan Jin, Sohyun Park et al.EMNLP 2024 · 5 citations
- Toward Annotator Group Bias in CrowdsourcingHaochen Liu, Joseph Thekinen, Sinem Mollaoglu, Da Tang et al.ACL 2022 · 20 citations
