Can Large Language Models Capture Dissenting Human Voices?
Noah Lee, Na An, James Thorne
摘要
Large language models (LLMs) have shown impressive achievements in solving a broad range of tasks. Augmented by instruction fine-tuning, LLMs have also been shown to generalize in zero-shot settings as well. However, whether LLMs closely align with the human disagreement distribution has not been well-studied, especially within the scope of natural language inference (NLI). In this paper, we evaluate the performance and alignment of LLM distribution with humans using two different techniques to estimate the multinomial distribution: Monte Carlo Estimation (MCE) and Log Probability Estimation (LPE). As a result, we show LLMs exhibit limited ability in solving NLI tasks and simultaneously fail to capture human disagreement distribution. The inference and human alignment performances plunge even further on data samples with high human disagreement levels, raising concerns about their natural language understanding (NLU) ability and their representativeness to a larger human population. 1 * Equal contribution 1 The source code for the experiments is available at https://github.com/xfactlab/emnlp2023-LLM-Disagreement . O Read the following and determine if the hypothesis can be inferred from the premise. Premise: She smiled back. Hypothesis: She was so happy she couldn't stop smiling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Which Demographics do LLMs Default to During Annotation?Johannes Schäfer, Aidan Combs, Christopher Bagdon, Jiahui Li 等ACL 2025 · 被引用 11 次
- How Hard is this Test Set? NLI Characterization by Exploiting Training DynamicsAdrian Cosma, Stefan Ruseti, Mihai Dascalu, Cornelia CarageaEMNLP 2024
- Voices in a Crowd: Searching for clusters of unique perspectivesNikolas Vitsakis, Amit Parekh, Ioannis KonstasEMNLP 2024
- Threading the Needle: Reweaving Chain-of-Thought Reasoning to Explain Human Label VariationBeiduo Chen, Yang Janet Liu, Anna Korhonen, Barbara PlankEMNLP 2025
- Label and Explanation Variation in LLM-Based Annotation: a Case Study in Natural Language InferenceArtur Kulmizev, Erika Lombart, Patrick Watrin, Marie-Catherine de MarneffeACL 2026
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson 等ICML 2023 · 被引用 908 次
- Cross-Task Generalization via Natural Language Crowdsourcing InstructionsSwaroop Mishra, Daniel Khashabi, Chitta Baral, Hannaneh HajishirziACL 2022 · 被引用 887 次
相关 Paper
- This is not a Dataset: A Large Negation Benchmark to Challenge Large Language ModelsIker García-Ferrero, Begoña Altuna, Javier Álvez, Itziar Gonzalez-Dios 等EMNLP 2023 · 被引用 8 次
- Large Language Models Do Multi-Label Classification DifferentlyMarcus Ma, Georgios Chochlakis, Niyantha Maruthu Pandiyan, Jesse Thomason 等EMNLP 2025
- Apathetic or Empathetic? Evaluating LLMs' Emotional Alignments with HumansJen-tse Huang, Man Ho Lam, Eric John Li, Shujie Ren 等NeurIPS 2024 · 被引用 63 次
- Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy CompetitionKehua Feng, Keyan Ding, Hongzhi Tan, Kede Ma 等ACL 2025
- Conditional and Modal Reasoning in Large Language ModelsWesley H. Holliday, Matthew Mandelkern, Cedegao ZhangEMNLP 2024 · 被引用 5 次
