Lune

EMNLP2025Top-tier venue

Learning Subjective Label Distributions via Sociocultural Descriptors

Mohammed Fayiz Parappan, Ricardo Henao

2025Year
5Citations

Abstract

Subjectivity in NLP tasks, e.g., toxicity classification, has emerged as a critical challenge precipitated by the increased deployment of NLP systems in content-sensitive domains. Conventional approaches aggregate annotator judgements (labels), ignoring minority perspectives and overlooking the influence of the sociocultural context behind such annotations. We propose a framework 1 where subjectivity in binary labels is modeled as an empirical distribution accounting for the variation in annotators through human values extracted from sociocultural descriptors using a language model. The framework also allows for downstream tasks such as population and sociocultural grouplevel majority label prediction. Experiments on three toxicity datasets covering human-chatbot conversations and social media posts annotated with diverse annotator pools demonstrate that our approach yields well-calibrated toxicity distribution predictions across binary toxicity labels, which are further used for majority label prediction across cultural subgroups, improving over existing methods.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 57cc65b4-5b2a-4661-bf09-e14d2670fc91

Builds on6

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines