Voices in a Crowd: Searching for clusters of unique perspectives
Nikolas Vitsakis, Amit Parekh, Ioannis Konstas
Abstract
Language models have been shown to reproduce underlying biases existing in their training data, which is the majority perspective by default.Proposed solutions aim to capture minority perspectives by either modelling annotator disagreements or grouping annotators based on shared metadata, both of which face significant challenges.We propose a framework that trains models without encoding annotator metadata, extracts latent embeddings informed by annotator behaviour, and creates clusters of similar opinions, that we refer to as voices.Resulting clusters are validated post-hoc via internal and external quantitative metrics, as well a qualitative analysis to identify the type of voice that each cluster represents.Our results demonstrate the strong generalisation capability of our framework, indicated by resulting clusters being adequately robust, while also capturing minority perspectives based on different demographic factors throughout two distinct datasets. 1Content Warning: This document contains and discusses examples of potentially offensive and toxic language.i) Disagreement-based (Metadata naive) MODEL per example (e.g., Ex. 1) Minority 0.4 Majority 0.6 ii) Metadata-based (Metadata info conditioned) MODEL L R L R per dataset (e.g., Ex. 1 & Ex .2) L L Disagreementbased Metadata constrained Captures dataset-level effects Dynamic grouping of annotators Number of identifiable voices Metadata-based Voices in a crowd 2 Any + metadata agnosticClimate change means the end of shopping. R LEco-towns could provide an inspiring blueprint for low-carbon living.Ex.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 133d5e47-d8f2-4dd6-80ae-a0ed352df36dCited by top-tier papers8
- Modeling Annotator Disagreement with Demographic-Aware Experts and Synthetic PerspectivesYinuo Xu, Veronica Derricks, Allison Earl, David JurgensACL 2026 · 8 citations
- Belief-Calibrated Multi-Agent Consensus Seeking for Complex NLP TasksWentao Deng, Jiahuan Pei, Zhiwei Xu, Zhaochun Ren et al.NeurIPS 2025 · 2 citations
- Value Profiles for Encoding Human VariationTaylor Sorensen, Pushkar Mishra, Roma Patel, Michael Henry Tessler et al.EMNLP 2025 · 2 citations
- NUTMEG: Separating Signal From Noise in Annotator DisagreementJonathan Ivey, Susan Gauch, David JurgensEMNLP 2025
- PERSEVAL: A Framework for Perspectivist Classification EvaluationSoda Marem Lo, Silvia Casola, Erhan Sezerer, Valerio Basile et al.EMNLP 2025
Builds on10
- FLAVA: A Foundational Language And Vision Alignment ModelAmanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon et al.CVPR 2022 · 483 citations
- Jury Learning: Integrating Dissenting Voices into Machine Learning ModelsMitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel et al.CHI 2022 · 134 citations
- Topic Discovery via Latent Space Clustering of Pretrained Language Model RepresentationsYu Meng, Yunyi Zhang, Jiaxin Huang, Yu Zhang et al.WWW 2022 · 73 citations
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 68 citations
- Isotropy in the Contextual Embedding Space: Clusters and ManifoldsXingyu Cai, Jiaji Huang, Yuchen Bian, Kenneth ChurchICLR 2021 · 50 citations
Related papers
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee et al.ICML 2023 · 764 citations
- Subjective Crowd Disagreements for Subjective Data: Uncovering Meaningful CrowdOpinion with Population-level LearningTharindu Cyril Weerasooriya, Sarah Luger, Saloni Poddar, Ashiqur R. KhudaBukhsh et al.ACL 2023 · 2 citations
- What Do Large Language Models Know About Opinions?Erfan Jahanparast, Zhiqing Hong, Serina ChangICLR 2026
- Investigating LLM-Powered Dissenting Minority Support in Power-Imbalanced Group Decision-Making: Counterargument and Mediation as Intervention StrategiesSoohwan Lee, Seoyeong Hwang, Mingyu Kim, Dajung Kim et al.CSCW 2026
- Large Language Models Develop Novel Social Biases Through Adaptive ExplorationAddison J. Wu, Ryan Liu, Xuechunzi Bai, Thomas GriffithsICML 2026 · 4 citations
