Lune

EMNLP2024顶会

Voices in a Crowd: Searching for clusters of unique perspectives

Nikolas Vitsakis, Amit Parekh, Ioannis Konstas

2024年份
8顶会引用

摘要

Language models have been shown to reproduce underlying biases existing in their training data, which is the majority perspective by default.Proposed solutions aim to capture minority perspectives by either modelling annotator disagreements or grouping annotators based on shared metadata, both of which face significant challenges.We propose a framework that trains models without encoding annotator metadata, extracts latent embeddings informed by annotator behaviour, and creates clusters of similar opinions, that we refer to as voices.Resulting clusters are validated post-hoc via internal and external quantitative metrics, as well a qualitative analysis to identify the type of voice that each cluster represents.Our results demonstrate the strong generalisation capability of our framework, indicated by resulting clusters being adequately robust, while also capturing minority perspectives based on different demographic factors throughout two distinct datasets. 1Content Warning: This document contains and discusses examples of potentially offensive and toxic language.i) Disagreement-based (Metadata naive) MODEL per example (e.g., Ex. 1) Minority 0.4 Majority 0.6 ii) Metadata-based (Metadata info conditioned) MODEL L R L R per dataset (e.g., Ex. 1 & Ex .2) L L Disagreementbased Metadata constrained Captures dataset-level effects Dynamic grouping of annotators Number of identifiable voices Metadata-based Voices in a crowd 2 Any + metadata agnosticClimate change means the end of shopping. R LEco-towns could provide an inspiring blueprint for low-carbon living.Ex.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 133d5e47-d8f2-4dd6-80ae-a0ed352df36d

引用它的顶会 Paper8

问问它们各自怎么用它

它引用的顶会 Paper10

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖