Lune

EMNLP2025顶会

Learning Subjective Label Distributions via Sociocultural Descriptors

Mohammed Fayiz Parappan, Ricardo Henao

2025年份
5被引次数

摘要

Subjectivity in NLP tasks, e.g., toxicity classification, has emerged as a critical challenge precipitated by the increased deployment of NLP systems in content-sensitive domains. Conventional approaches aggregate annotator judgements (labels), ignoring minority perspectives and overlooking the influence of the sociocultural context behind such annotations. We propose a framework 1 where subjectivity in binary labels is modeled as an empirical distribution accounting for the variation in annotators through human values extracted from sociocultural descriptors using a language model. The framework also allows for downstream tasks such as population and sociocultural grouplevel majority label prediction. Experiments on three toxicity datasets covering human-chatbot conversations and social media posts annotated with diverse annotator pools demonstrate that our approach yields well-calibrated toxicity distribution predictions across binary toxicity labels, which are further used for majority label prediction across cultural subgroups, improving over existing methods.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 57cc65b4-5b2a-4661-bf09-e14d2670fc91

它引用的顶会 Paper6

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖