Voices in a Crowd: Searching for clusters of unique perspectives
Nikolas Vitsakis, Amit Parekh, Ioannis Konstas
摘要
Language models have been shown to reproduce underlying biases existing in their training data, which is the majority perspective by default.Proposed solutions aim to capture minority perspectives by either modelling annotator disagreements or grouping annotators based on shared metadata, both of which face significant challenges.We propose a framework that trains models without encoding annotator metadata, extracts latent embeddings informed by annotator behaviour, and creates clusters of similar opinions, that we refer to as voices.Resulting clusters are validated post-hoc via internal and external quantitative metrics, as well a qualitative analysis to identify the type of voice that each cluster represents.Our results demonstrate the strong generalisation capability of our framework, indicated by resulting clusters being adequately robust, while also capturing minority perspectives based on different demographic factors throughout two distinct datasets. 1Content Warning: This document contains and discusses examples of potentially offensive and toxic language.i) Disagreement-based (Metadata naive) MODEL per example (e.g., Ex. 1) Minority 0.4 Majority 0.6 ii) Metadata-based (Metadata info conditioned) MODEL L R L R per dataset (e.g., Ex. 1 & Ex .2) L L Disagreementbased Metadata constrained Captures dataset-level effects Dynamic grouping of annotators Number of identifiable voices Metadata-based Voices in a crowd 2 Any + metadata agnosticClimate change means the end of shopping. R LEco-towns could provide an inspiring blueprint for low-carbon living.Ex.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Modeling Annotator Disagreement with Demographic-Aware Experts and Synthetic PerspectivesYinuo Xu, Veronica Derricks, Allison Earl, David JurgensACL 2026 · 被引用 8 次
- Belief-Calibrated Multi-Agent Consensus Seeking for Complex NLP TasksWentao Deng, Jiahuan Pei, Zhiwei Xu, Zhaochun Ren 等NeurIPS 2025 · 被引用 2 次
- Value Profiles for Encoding Human VariationTaylor Sorensen, Pushkar Mishra, Roma Patel, Michael Henry Tessler 等EMNLP 2025 · 被引用 2 次
- NUTMEG: Separating Signal From Noise in Annotator DisagreementJonathan Ivey, Susan Gauch, David JurgensEMNLP 2025
- PERSEVAL: A Framework for Perspectivist Classification EvaluationSoda Marem Lo, Silvia Casola, Erhan Sezerer, Valerio Basile 等EMNLP 2025
它引用的顶会 Paper10
- FLAVA: A Foundational Language And Vision Alignment ModelAmanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon 等CVPR 2022 · 被引用 483 次
- Jury Learning: Integrating Dissenting Voices into Machine Learning ModelsMitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel 等CHI 2022 · 被引用 134 次
- Topic Discovery via Latent Space Clustering of Pretrained Language Model RepresentationsYu Meng, Yunyi Zhang, Jiaxin Huang, Yu Zhang 等WWW 2022 · 被引用 73 次
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 被引用 68 次
- Isotropy in the Contextual Embedding Space: Clusters and ManifoldsXingyu Cai, Jiaji Huang, Yuchen Bian, Kenneth ChurchICLR 2021 · 被引用 50 次
相关 Paper
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee 等ICML 2023 · 被引用 764 次
- Subjective Crowd Disagreements for Subjective Data: Uncovering Meaningful CrowdOpinion with Population-level LearningTharindu Cyril Weerasooriya, Sarah Luger, Saloni Poddar, Ashiqur R. KhudaBukhsh 等ACL 2023 · 被引用 2 次
- What Do Large Language Models Know About Opinions?Erfan Jahanparast, Zhiqing Hong, Serina ChangICLR 2026
- Investigating LLM-Powered Dissenting Minority Support in Power-Imbalanced Group Decision-Making: Counterargument and Mediation as Intervention StrategiesSoohwan Lee, Seoyeong Hwang, Mingyu Kim, Dajung Kim 等CSCW 2026
- Large Language Models Develop Novel Social Biases Through Adaptive ExplorationAddison J. Wu, Ryan Liu, Xuechunzi Bai, Thomas GriffithsICML 2026 · 被引用 4 次
