Evaluating Language Model Pluralism through In-the-wild Crowd Discussions
Gagan Mundada, Rohan Surana, Nandhini Swaminathan, Bodhisattwa Prasad Majumder, Junda Wu, Julian J. McAuley, Zhouhang Xie
Abstract
When answering subjective questions, an ideal LLM should surface diverse plausible perspectives rather than favoring a single viewpoint, a characteristic known as pluralism. Recent studies show that modern LLMs optimized through preference alignment systematically favor certain positions on subjective queries, making pluralism evaluation increasingly important. However, existing evaluation methods focus dominantly on multiple-choice and question-answering tasks, leaving open-ended generation largely unaddressed. We propose PLURALEVAL, an evaluation framework that assesses LLM pluralism in open-ended generation by comparing outputs against free-form crowd responses. Our approach decomposes ground-truth responses into atomic, non-overlapping claims, then evaluates whether LLMs adequately cover this diverse claim space. We then introduce WILD-SCOPE, a multi-domain dataset of natural crowd responses, and demonstrate that PLU-RALEVAL captures novel insights, such as the collapse of pluralism through sycophancy, where LLM systematically degrades in Overton pluralism when a user's belief is revealed. Finally, we discuss the value and actionable insights for preserving and encouraging pluralism from LLM deployers' side 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a1916ee3-49fc-4eea-bbef-06d8583e8d7dBuilds on20
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch et al.ICLR 2021 · 878 citations
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent DebateChi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu et al.ICLR 2024 · 871 citations
- Towards Understanding Sycophancy in Language ModelsMrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud et al.ICLR 2024 · 762 citations
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun et al.NeurIPS 2021 · 606 citations
Related papers
- Benchmarking Overton Pluralism in LLMsElinor Poole-Dayan, Jiayi Wu, Taylor Sorensen, Jiaxin Pei et al.ICLR 2026 · 9 citations
- Modular Pluralism: Pluralistic Alignment via Multi-LLM CollaborationShangbin Feng, Taylor Sorensen, Yuhan Liu, Jillian Fisher et al.EMNLP 2024 · 12 citations
- VITAL: A New Dataset for Benchmarking Pluralistic Alignment in HealthcareAnudeex Shetty, Amin Beheshti, Mark Dras, Usman NaseemACL 2025 · 13 citations
- PerSpectra: A Scalable and Configurable Pluralist Benchmark of Perspectives from ArgumentsShangrui Nie, Kian Omoomi, Lucie Flek, Zhixue Zhao et al.ICLR 2026 · 3 citations
- Learning Personalized Alignment for Evaluating Open-ended Text GenerationDanqing Wang, Kevin Yang, Hanlin Zhu, Xiaomeng Yang et al.EMNLP 2024 · 2 citations
