Unmasking and Improving Data Credibility: A Study with Datasets for Training Harmless Language Models
Zhaowei Zhu, Jialu Wang, Hao Cheng, Yang Liu
Abstract
Language models have shown promise in various tasks but can be affected by undesired data during training, fine-tuning, or alignment. For example, if some unsafe conversations are wrongly annotated as safe ones, the model fine-tuned on these samples may be harmful. Therefore, the correctness of annotations, i.e., the credibility of the dataset, is important. This study focuses on the credibility of real-world datasets, including the popular benchmarks Jigsaw Civil Comments, Anthropic Harmless & Red Team, PKU BeaverTails & SafeRLHF, that can be used for training a harmless language model. Given the cost and difficulty of cleaning these datasets by humans, we introduce a systematic framework for evaluating the credibility of datasets, identifying label errors, and evaluating the influence of noisy labels in the curated language data, specifically focusing on unsafe comments and conversation classification. With the framework, we find and fix an average of 6.16% label errors in 11 datasets constructed from the above benchmarks. The data credibility and downstream learning performance can be remarkably improved by directly fixing label errors, indicating the significance of cleaning existing real-world datasets. We provide an open-source tool, Docta, for data cleaning at https://github.com/Docta-ai/docta .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fcfc9394-253f-4d52-b7af-b88700b11990Cited by top-tier papers8
- EXP-Bench: Can AI Conduct AI Research Experiments?Patrick Tser Jern Kon, Qiuyi Ding, Jiachen Liu, Xinyi Zhu et al.ICLR 2026 · 35 citations
- FedFixer: Mitigating Heterogeneous Label Noise in Federated LearningXinyuan Ji, Zhaowei Zhu, Wei Xi, Olga Gadyatskaya et al.AAAI 2024 · 30 citations
- Robust Preference Alignment via Directional Neighborhood ConsensusRuochen Mao, Yuling Shi, Xiaodong Gu, Jiaheng WeiICLR 2026 · 2 citations
- Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark StudyDongGeon Lee, Joonwon Jang, Jihae Jeong, Hwanjo YuEMNLP 2025 · 1 citation
- LLM Unlearning via Loss Adjustment with Only Forget DataYaxuan Wang, Jiaheng Wei, Chris Yuhao Liu, Jinlong Pang et al.ICLR 2025
Builds on32
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
Related papers
- Are we building on the rock? on the importance of data preprocessing for code summarizationLin Shi, Fangwen Mu, Xiao Chen, Song Wang et al.FSE 2022 · 65 citations
- Noisy but Valid: Robust Statistical Evaluation of LLMs with Imperfect JudgesChen Feng, Minghe Shen, Ananth Balashankar, Carsten Gerner-Beuerle et al.ICLR 2026 · 24 citations
- SafeConv: Explaining and Correcting Conversational Unsafe BehaviorMian Zhang, Lifeng Jin, Linfeng Song, Haitao Mi et al.ACL 2023 · 5 citations
- UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated ImagesYiting Qu, Xinyue Shen, Yixin Wu, Michael Backes et al.CCS 2025 · 1 citation
- Human Uncertainty-Aware Data Selection and Automatic Labeling in Visual Question AnsweringJian Lan, Zhicheng Liu, Udo Schlegel, Raoyuan Zhao et al.ICLR 2026 · 2 citations
