Learning Subjective Label Distributions via Sociocultural Descriptors
Mohammed Fayiz Parappan, Ricardo Henao
摘要
Subjectivity in NLP tasks, e.g., toxicity classification, has emerged as a critical challenge precipitated by the increased deployment of NLP systems in content-sensitive domains. Conventional approaches aggregate annotator judgements (labels), ignoring minority perspectives and overlooking the influence of the sociocultural context behind such annotations. We propose a framework 1 where subjectivity in binary labels is modeled as an empirical distribution accounting for the variation in annotators through human values extracted from sociocultural descriptors using a language model. The framework also allows for downstream tasks such as population and sociocultural grouplevel majority label prediction. Experiments on three toxicity datasets covering human-chatbot conversations and social media posts annotated with diverse annotator pools demonstrate that our approach yields well-calibrated toxicity distribution predictions across binary toxicity labels, which are further used for majority label prediction across cultural subgroups, improving over existing methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Jury Learning: Integrating Dissenting Voices into Machine Learning ModelsMitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel 等CHI 2022 · 被引用 134 次
- The Disagreement Deconvolution: Bringing Machine Learning Performance Metrics In Line With RealityMitchell L. Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto 等CHI 2021 · 被引用 100 次
- When the Majority is Wrong: Modeling Annotator Disagreement for Subjective TasksEve Fleisig, Rediet Abebe, Dan KleinEMNLP 2023 · 被引用 11 次
- Soft-Label Integration for Robust Toxicity ClassificationZelei Cheng, Xian Wu, Jiahao Yu, Shuo Han 等NeurIPS 2024 · 被引用 7 次
- How Far Can We Extract Diverse Perspectives from Large Language Models?Shirley Anugrah Hayati, Minhwa Lee, Dheeraj Rajagopal, Dongyeop KangEMNLP 2024 · 被引用 6 次
相关 Paper
- PERSEVAL: A Framework for Perspectivist Classification EvaluationSoda Marem Lo, Silvia Casola, Erhan Sezerer, Valerio Basile 等EMNLP 2025
- ModelCitizens: Representing Community Voices in Online SafetyAshima Suvarna, Christina Chance, Karolina Naranjo, Hamid Palangi 等EMNLP 2025
- Same Same, But Different: Conditional Multi-Task Learning for Demographic-Specific Toxicity DetectionSoumyajit Gupta, Sooyong Lee, Maria De-Arteaga, Matthew LeaseWWW 2023 · 被引用 17 次
- MultiPICo: Multilingual Perspectivist Irony CorpusSilvia Casola, Simona Frenda, Soda Marem Lo, Erhan Sezerer 等ACL 2024 · 被引用 2 次
- Towards Author-informed NLP: Mind the Social BiasInbar Pendzel, Einat MinkovEMNLP 2025
