Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model Performance
Omer Nahum, Nitay Calderon, Orgad Keller, Idan Szpektor, Roi Reichart
Abstract
NLP benchmarks rely on standardized datasets for training and evaluating models and are crucial for advancing the field. Traditionally, expert annotations ensure high-quality labels; however, the cost of expert annotation does not scale well with the growing demand for larger datasets required by modern models. While crowd-sourcing provides a more scalable solution, it often comes at the expense of annotation precision and consistency. Recent advancements in large language models (LLMs) offer new opportunities to enhance the annotation process, particularly for detecting label errors in existing datasets. In this work, we consider the recent approach of LLM-as-a-judge, leveraging an ensemble of LLMs to flag potentially mislabeled examples. We conduct a case study on four factual consistency datasets from the TRUE benchmark, spanning diverse NLP tasks, and on SummEval, which uses Likertscale ratings of summary quality across multiple dimensions. We empirically analyze the labeling quality of existing datasets and compare expert, crowd-sourced, and LLM-based annotations in terms of the agreement, label quality, and efficiency, demonstrating the strengths and limitations of each annotation method. Our findings reveal a substantial number of label errors, which, when corrected, induce a significant upward shift in reported model performance. This suggests that many of the LLMs' so-called mistakes are due to label errors rather than genuine model failures. Additionally, we discuss the implications of mislabeled data and propose methods to mitigate them in training to improve performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 30d716b5-c77d-458c-8430-b53de0f7903dCited by top-tier papers8
- Prompting in the Dark: Assessing Human Performance in Prompt Engineering for Data Labeling When Gold Labels Are AbsentZeyu He, Saniya Naphade, Ting-Hao 'Kenneth' HuangCHI 2025 · 22 citations
- Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric FactualityNitay Calderon, Eyal Ben-David, Zorik Gekhman, Eran Ofek et al.ICML 2026 · 8 citations
- Empowering Tabular Data Preparation with Language Models: Why and How?Mengshi Chen, Yuxiang Sun, Tengchao Li, Jianwei Wang et al.ACL 2026 · 4 citations
- DeepFact: Co-Evolving Benchmarks and Agents for Deep Research FactualityYukun Huang, Leonardo F. R. Ribeiro, Momchil Hardalov, Bhuwan Dhingra et al.ACL 2026 · 2 citations
- BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice BenchmarksNishant Balepur, Bhavya Rajasekaran, Hyunjin Jane Oh, Michael Xie et al.ACL 2026 · 1 citation
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
Related papers
- SummEdits: Measuring LLM Ability at Factual Reasoning Through The Lens of SummarizationPhilippe Laban, Wojciech Kryscinski, Divyansh Agarwal, Alexander R. Fabbri et al.EMNLP 2023 · 29 citations
- Limits to scalable evaluation at the frontier: LLM as judge won't beat twice the dataFlorian E. Dorner, Vivian Yvonne Nastl, Moritz HardtICLR 2025
- STORYSUMM: Evaluating Faithfulness in Story SummarizationMelanie Subbiah, Faisal Ladhak, Akankshya Mishra, Griffin Adams et al.EMNLP 2024 · 2 citations
- SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality EvaluationHui Wang, Jinghua Zhao, Yifan Yang, Shujie Liu et al.ACL 2026 · 21 citations
- TrueTeacher: Learning Factual Consistency Evaluation with Large Language ModelsZorik Gekhman, Jonathan Herzig, Roee Aharoni, Chen Elkind et al.EMNLP 2023 · 24 citations
