What Can We Learn from Collective Human Opinions on Natural Language Inference Data?
Yixin Nie, Xiang Zhou, Mohit Bansal
Abstract
Despite the subjective nature of many NLP tasks, most NLU evaluations have focused on using the majority label with presumably high agreement as the ground truth. Less attention has been paid to the distribution of human opinions. We collect ChaosNLI, a dataset with a total of 464,500 annotations to study Collective HumAn OpinionS in oft-used NLI evaluation sets. This dataset is created by collecting 100 annotations per example for 3,113 examples in SNLI and MNLI and 1,532 examples in αNLI. Analysis reveals that: (1) high human disagreement exists in a noticeable amount of examples in these datasets; (2) the state-of-the-art models lack the ability to recover the distribution over human labels; (3) models achieve near-perfect accuracy on the subset of data with a high level of human agreement, whereas they can barely beat a random guess on the data with low levels of human agreement, which compose most of the common errors made by state-of-the-art models on the evaluation sets. This questions the validity of improving model performance on old metrics for the low-agreement part of evaluation datasets. Hence, we argue for a detailed examination of human agreement in future data collection efforts, and evaluating model outputs against the distribution over collective human opinions. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9192ef34-51d7-454a-8ab4-82cbc6d60dbfCited by top-tier papers37
- SimBench: Benchmarking the Ability of Large Language Models to Simulate Human BehaviorsTiancheng Hu, Joachim Baumann, Lorenzo Lupo, Nigel Collier et al.ICLR 2026 · 61 citations
- Validating LLM-as-a-Judge Systems under Rating IndeterminacyLuke Guerdan, Solon Barocas, Kenneth Holstein, Hanna M. Wallach et al.NeurIPS 2025 · 31 citations
- Conformalized Credal Set PredictorsAlireza Javanmardi, David Stutz, Eyke HüllermeierNeurIPS 2024 · 28 citations
- Zero and Few-shot Semantic Parsing with Ambiguous InputsElias Stengel-Eskin, Kyle Rawlins, Benjamin Van DurmeICLR 2024 · 27 citations
- MiniCheck: Efficient Fact-Checking of LLMs on Grounding DocumentsLiyan Tang, Philippe Laban, Greg DurrettEMNLP 2024 · 26 citations
Builds on4
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- Abductive Commonsense ReasoningChandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi et al.ICLR 2020 · 521 citations
Related papers
- Detecting Label Errors by Using Pre-Trained Language ModelsDerek Chong, Jenny Hong, Christopher D. ManningEMNLP 2022 · 8 citations
- Forest vs Tree: The (N, K) Trade-off in Reproducible ML EvaluationDeepak Pandita, Flip Korn, Chris Welty, Christopher M. HomanAAAI 2026 · 2 citations
- IndoNLI: A Natural Language Inference Dataset for IndonesianRahmad Mahendra, Alham Fikri Aji, Samuel Louvan, Fahrurrozi Rahman et al.EMNLP 2021 · 14 citations
- Can Large Language Models Capture Dissenting Human Voices?Noah Lee, Na An, James ThorneEMNLP 2023 · 8 citations
- What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks?Nikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt et al.ACL 2021
