Detecting Label Errors by Using Pre-Trained Language Models
Derek Chong, Jenny Hong, Christopher D. Manning
摘要
We show that large pre-trained language models are inherently highly capable of identifying label errors in natural language datasets: simply examining out-of-sample data points in descending order of fine-tuned task loss significantly outperforms more complex error-detection mechanisms proposed in previous work. To this end, we contribute a novel method for introducing realistic, human-originated label noise into existing crowdsourced datasets such as SNLI and TweetNLP. We show that this noise has similar properties to real, hand-verified label errors, and is harder to detect than existing synthetic noise, creating challenges for model robustness.We argue that human-originated noise is a better standard for evaluation than synthetic noise. Finally, we use crowdsourced verification to evaluate the detection of real errors on IMDB, Amazon Reviews, and Recon, and confirm that pre-trained models perform at a 9–36% higher absolute Area Under the Precision-Recall Curve than existing models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Unmasking and Improving Data Credibility: A Study with Datasets for Training Harmless Language ModelsZhaowei Zhu, Jialu Wang, Hao Cheng, Yang LiuICLR 2024 · 被引用 30 次
- Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model PerformanceOmer Nahum, Nitay Calderon, Orgad Keller, Idan Szpektor 等EMNLP 2025 · 被引用 9 次
- NoiseBench: Benchmarking the Impact of Real Label Noise on Named Entity RecognitionElena Merdjanovska, Ansar Aynetdinov, Alan AkbikEMNLP 2024 · 被引用 5 次
- Rethinking Evaluation for LLM Hallucination Detection: A Desiderata, A New RAG-based Benchmark, New InsightsWenbo Chen, Veena Padmanabhan, Tootiya Giyahchi, Elaine Wong 等ACL 2026 · 被引用 1 次
- Model Editing as a Robust and Denoised variant of DPO: A Case Study on ToxicityRheeya Uppaal, Apratim Dey, Yiting He, Yiqiao Zhong 等ICLR 2025
它引用的顶会 Paper7
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Normalized Loss Functions for Deep Learning with Noisy LabelsXingjun Ma, Hanxun Huang, Yisen Wang, Simone Romano 等ICML 2020 · 被引用 547 次
- Learning with Noisy Labels Revisited: A Study Using Real-World Human AnnotationsJiaheng Wei, Zhaowei Zhu, Hao Cheng, Tongliang Liu 等ICLR 2022 · 被引用 338 次
- Peer Loss Functions: Learning from Noisy Labels without Knowing Noise RatesYang Liu, Hongyi GuoICML 2020 · 被引用 280 次
- FINE Samples for Learning with Noisy LabelsTaehyeon Kim, Jongwoo Ko, Sangwook Cho, Jinhwan Choi 等NeurIPS 2021 · 被引用 145 次
相关 Paper
- LEMoN: Label Error Detection using Multimodal NeighborsHaoran Zhang, Aparna Balagopalan, Nassim Oufattole, Hyewon Jeong 等ICML 2025
- VariErr NLI: Separating Annotation Error from Human Label VariationLeon Weber-Genzel, Siyao Peng, Marie-Catherine de Marneffe, Barbara PlankACL 2024
- New Protocols and Negative Results for Textual Entailment Data CollectionSamuel R. Bowman, Jennimaria Palomaki, Livio Baldini Soares, Emily PitlerEMNLP 2020 · 被引用 3 次
- Understanding and Mitigating the Label Noise in Pre-training on Downstream TasksHao Chen, Jindong Wang, Ankit Shah, Ran Tao 等ICLR 2024 · 被引用 49 次
- Analysing the Noise Model Error for Realistic Noisy Label DataMichael A. Hedderich, Dawei Zhu, Dietrich KlakowAAAI 2021 · 被引用 26 次
