Label and Explanation Variation in LLM-Based Annotation: a Case Study in Natural Language Inference
Artur Kulmizev, Erika Lombart, Patrick Watrin, Marie-Catherine de Marneffe
摘要
Large language models (LLMs) have shown considerable promise for annotation purposes, yet questions remain about their ability to capture human label variation (HLV) -genuine disagreement between annotators often observed across NLP tasks. Here, we investigate how label and explanation variation manifests within and across LLMs with respect to the Natural Language Inference (NLI) task in English. Using zero-shot prompting with exact human annotation instructions, we treat individual model generations as participants and examine three response sampling strategies: varying generation parameters, leveraging within-family model size differences, and pooling responses from distinct LLMs. We show that, while model ensembles can generate label distributions similar to humans, they likewise exhibit distinct, idiosyncratic judgments and disagreement patterns. We further analyze explanation variation, observing that, although models generate longer explanations than humans, they demonstrate substantially less stylistic diversity. Our findings suggest that, while LLMs may serve as useful tools for generating diverse annotations, they should not be viewed as drop-in replacements for human annotatorsparticularly in applications requiring authentic representation of diversity in human judgments, such as NLI.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team PerformanceGagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok 等CHI 2021 · 被引用 713 次
- Generating Training Data with Language Models: Towards Zero-Shot Language UnderstandingYu Meng, Jiaxin Huang, Yu Zhang, Jiawei HanNeurIPS 2022 · 被引用 309 次
- Human-LLM Collaborative Annotation Through Effective Verification of LLM LabelsXinru Wang, Hannah Kim, Sajjadur Rahman, Kushan Mitra 等CHI 2024 · 被引用 127 次
- Large Language Models for Data Annotation and Synthesis: A SurveyZhen Tan, Dawei Li, Song Wang, Alimohammad Beigi 等EMNLP 2024 · 被引用 119 次
相关 Paper
- LiTEx: A Linguistic Taxonomy of Explanations for Understanding Within-Label Variation in Natural Language InferencePingjun Hong, Beiduo Chen, Siyao Peng, Marie-Catherine de Marneffe 等EMNLP 2025 · 被引用 1 次
- Can Large Language Models Capture Dissenting Human Voices?Noah Lee, Na An, James ThorneEMNLP 2023 · 被引用 8 次
- Estimating LLM Consistency: A User Baseline vs Surrogate MetricsXiaoyuan Wu, Weiran Lin, Omer Akgul, Lujo BauerEMNLP 2025
- Improving Diversity of Demographic Representation in Large Language Models via Collective-Critiques and Self-VotingPreethi Lahoti, Nicholas Blumm, Xiao Ma, Raghavendra Kotikalapudi 等EMNLP 2023 · 被引用 14 次
- Quantifying the Persona Effect in LLM SimulationsTiancheng Hu, Nigel CollierACL 2024 · 被引用 22 次
