Lune

ACL2026Top-tier venue

Label and Explanation Variation in LLM-Based Annotation: a Case Study in Natural Language Inference

Artur Kulmizev, Erika Lombart, Patrick Watrin, Marie-Catherine de Marneffe

2026Year

Abstract

Large language models (LLMs) have shown considerable promise for annotation purposes, yet questions remain about their ability to capture human label variation (HLV) -genuine disagreement between annotators often observed across NLP tasks. Here, we investigate how label and explanation variation manifests within and across LLMs with respect to the Natural Language Inference (NLI) task in English. Using zero-shot prompting with exact human annotation instructions, we treat individual model generations as participants and examine three response sampling strategies: varying generation parameters, leveraging within-family model size differences, and pooling responses from distinct LLMs. We show that, while model ensembles can generate label distributions similar to humans, they likewise exhibit distinct, idiosyncratic judgments and disagreement patterns. We further analyze explanation variation, observing that, although models generate longer explanations than humans, they demonstrate substantially less stylistic diversity. Our findings suggest that, while LLMs may serve as useful tools for generating diverse annotations, they should not be viewed as drop-in replacements for human annotatorsparticularly in applications requiring authentic representation of diversity in human judgments, such as NLI.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 178637a9-3beb-4646-b717-8474a3661d8b

Builds on13

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines