Noise Correction on Subjective Datasets
Uthman Jinadu, Yi Ding
Abstract
Incorporating every annotator's perspective is crucial for unbiased data modeling. Annotator fatigue and changing opinions over time can distort dataset annotations. To combat this, we propose to learn a more accurate representation of diverse opinions by utilizing multitask learning in conjunction with loss-based label correction. We show that using our novel formulation, we can cleanly separate agreeing and disagreeing annotations. Furthermore, this method provides a controllable way to encourage or discourage disagreement. We demonstrate that this modification can improve prediction performance in a single or multi-annotator setting. Lastly, we show that this method remains robust to additional label noise that is applied to subjective data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on6
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 1,326 citations
- Early-Learning Regularization Prevents Memorization of Noisy LabelsSheng Liu, Jonathan Niles-Weed, Narges Razavian, Carlos Fernandez-GrandaNeurIPS 2020 · 798 citations
- Impact of Annotator Demographics on Sentiment Dataset LabelingYi Ding, Jacob You, Tonja-Katrin Machulla, Jennifer Jacobs et al.CSCW 2022 · 20 citations
- GoEmotions: A Dataset of Fine-Grained EmotionsDorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan S. Cowen et al.ACL 2020 · 16 citations
- Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators' DisagreementElisa Leonardelli, Stefano Menini, Alessio Palmero Aprosio, Marco Guerini et al.EMNLP 2021 · 2 citations
Related papers
- QuMAB: Query-based Multi-annotator Behavior Pattern LearningLiyun Zhang, Zheng Lian, Hong Liu, Takanori Takebe et al.AAAI 2026 · 3 citations
- Aggregating Complex Annotations via Merging and MatchingAlexander Braylan, Matthew LeaseKDD 2021 · 8 citations
- Collaborative Refining for Learning from Inaccurate LabelsBin Han, Yi-Xuan Sun, Ya-Lin Zhang, Libang Zhang et al.NeurIPS 2024 · 4 citations
- Discrepancy Ratio: Evaluating Model Performance When Even Experts Disagree on the TruthIgor Lovchinsky, Alon Daks, Israel Malkin, Pouya Samangouei et al.ICLR 2020 · 11 citations
- Learning Calibrated Medical Image Segmentation via Multi-Rater Agreement ModelingWei Ji, Shuang Yu, Junde Wu, Kai Ma et al.CVPR 2021
