Toward Annotator Group Bias in Crowdsourcing
Haochen Liu, Joseph Thekinen, Sinem Mollaoglu, Da Tang, Ji Yang, Youlong Cheng, Hui Liu, Jiliang Tang
Abstract
Crowdsourcing has emerged as a popular approach for collecting annotated data to train supervised machine learning models. However, annotator bias can lead to defective annotations. Though there are a few works investigating individual annotator bias, the group effects in annotators are largely overlooked. In this work, we reveal that annotators within the same demographic group tend to show consistent group bias in annotation tasks and thus we conduct an initial study on annotator group bias. We first empirically verify the existence of annotator group bias in various real-world crowdsourcing datasets. Then, we develop a novel probabilistic graphical framework GroupAnno to capture annotator group bias with a new extended Expectation Maximization (EM) training algorithm. We conduct experiments on both synthetic and real-world datasets. Experimental results demonstrate the effectiveness of our model in modeling annotator group bias in label aggregation and model learning over competitive baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Generating Data to Mitigate Spurious Correlations in Natural Language Inference DatasetsYuxiang Wu, Matt Gardner, Pontus Stenetorp, Pradeep DasigiACL 2022 · 74 citations
- Uncovering and Quantifying Social Biases in Code GenerationYan Liu, Xiaokang Chen, Yan Gao, Zhe Su et al.NeurIPS 2023 · 47 citations
- The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Language ModelsYan Liu, Yu Liu, Xiaokang Chen, Pin-Yu Chen et al.ICLR 2024 · 32 citations
- Verification and Refinement of Natural Language Explanations through LLM-Symbolic Theorem ProvingXin Quan, Marco Valentino, Louise A. Dennis, André FreitasEMNLP 2024 · 9 citations
- Explanation-based Finetuning Makes Models More Robust to Spurious CuesJosh Magnus Ludan, Yixuan Meng, Tai Nguyen, Saurabh Shah et al.ACL 2023 · 5 citations
Builds on1
Related papers
- Impact of Annotator Demographics on Sentiment Dataset LabelingYi Ding, Jacob You, Tonja-Katrin Machulla, Jennifer Jacobs et al.CSCW 2022 · 20 citations
- CrowdGP: a Gaussian Process Model for Inferring Relevance from Crowd AnnotationsDan Li, Zhaochun Ren, Evangelos KanoulasWWW 2021 · 6 citations
- Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text PerceptionsMatthias Orlikowski, Jiaxin Pei, Paul Röttger, Philipp Cimiano et al.ACL 2025 · 34 citations
- Optimal Fair Aggregation of Crowdsourced Noisy Labels using Demographic Parity ConstraintsGabriel Singer, Samuel Gruffaz, Olivier VO VAN, Nicolas Vayatis et al.ICML 2026
- When the Majority is Wrong: Modeling Annotator Disagreement for Subjective TasksEve Fleisig, Rediet Abebe, Dan KleinEMNLP 2023 · 11 citations
