Unmasking and Improving Data Credibility: A Study with Datasets for Training Harmless Language Models
Zhaowei Zhu, Jialu Wang, Hao Cheng, Yang Liu
摘要
Language models have shown promise in various tasks but can be affected by undesired data during training, fine-tuning, or alignment. For example, if some unsafe conversations are wrongly annotated as safe ones, the model fine-tuned on these samples may be harmful. Therefore, the correctness of annotations, i.e., the credibility of the dataset, is important. This study focuses on the credibility of real-world datasets, including the popular benchmarks Jigsaw Civil Comments, Anthropic Harmless & Red Team, PKU BeaverTails & SafeRLHF, that can be used for training a harmless language model. Given the cost and difficulty of cleaning these datasets by humans, we introduce a systematic framework for evaluating the credibility of datasets, identifying label errors, and evaluating the influence of noisy labels in the curated language data, specifically focusing on unsafe comments and conversation classification. With the framework, we find and fix an average of 6.16% label errors in 11 datasets constructed from the above benchmarks. The data credibility and downstream learning performance can be remarkably improved by directly fixing label errors, indicating the significance of cleaning existing real-world datasets. We provide an open-source tool, Docta, for data cleaning at https://github.com/Docta-ai/docta .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- EXP-Bench: Can AI Conduct AI Research Experiments?Patrick Tser Jern Kon, Qiuyi Ding, Jiachen Liu, Xinyi Zhu 等ICLR 2026 · 被引用 35 次
- FedFixer: Mitigating Heterogeneous Label Noise in Federated LearningXinyuan Ji, Zhaowei Zhu, Wei Xi, Olga Gadyatskaya 等AAAI 2024 · 被引用 30 次
- Robust Preference Alignment via Directional Neighborhood ConsensusRuochen Mao, Yuling Shi, Xiaodong Gu, Jiaheng WeiICLR 2026 · 被引用 2 次
- Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark StudyDongGeon Lee, Joonwon Jang, Jihae Jeong, Hwanjo YuEMNLP 2025 · 被引用 1 次
- LLM Unlearning via Loss Adjustment with Only Forget DataYaxuan Wang, Jiaheng Wei, Chris Yuhao Liu, Jinlong Pang 等ICLR 2025
它引用的顶会 Paper32
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
相关 Paper
- Are we building on the rock? on the importance of data preprocessing for code summarizationLin Shi, Fangwen Mu, Xiao Chen, Song Wang 等FSE 2022 · 被引用 65 次
- Noisy but Valid: Robust Statistical Evaluation of LLMs with Imperfect JudgesChen Feng, Minghe Shen, Ananth Balashankar, Carsten Gerner-Beuerle 等ICLR 2026 · 被引用 24 次
- SafeConv: Explaining and Correcting Conversational Unsafe BehaviorMian Zhang, Lifeng Jin, Linfeng Song, Haitao Mi 等ACL 2023 · 被引用 5 次
- UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated ImagesYiting Qu, Xinyue Shen, Yixin Wu, Michael Backes 等CCS 2025 · 被引用 1 次
- Human Uncertainty-Aware Data Selection and Automatic Labeling in Visual Question AnsweringJian Lan, Zhicheng Liu, Udo Schlegel, Raoyuan Zhao 等ICLR 2026 · 被引用 2 次
