Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks
Hao Chen, Jindong Wang, Ankit Shah, Ran Tao, Hongxin Wei, Xing Xie, Masashi Sugiyama, Bhiksha Raj
Abstract
Pre-training on large-scale datasets and then fine-tuning on downstream tasks have become a standard practice in deep learning. However, pre-training datasets, while inaccessible or too expensive to handle, often contain label noise that may adversely affect the generalization of the model and pose unexpected risks. This paper aims to understand the nature of noise in pre-training datasets and then mitigate its impact on downstream tasks. Specifically, through extensive experiments of supervised pre-training models on synthetic noisy ImageNet-1K and YFCC15M datasets, we demonstrate that while slight noise in pre-training can benefit in-domain (ID) performance, where the training and testing data share the same distribution, it always deteriorates out-of-domain (OOD) performance, where training and testing distributions are different. We empirically ascertain that the reason behind is noise in pre-training shapes the feature space differently. We then propose a light-weight black-box tuning method (NMTune) to affine the feature space to mitigate the malignant effect of noise and improve generalization on both ID and OOD tasks, considering that one may not be able to access or fully fine-tune the pre-trained models. We conduct extensive experiments on popular vision and language models including APIs that are supervised and self-supervised pre-trained on real data for evaluation. Our results show the importance of this novel and fundamental research direction, which we term Noisy Model Learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b08e972e-a2b5-4af7-9992-dba1132645a5Cited by top-tier papers9
- Slight Corruption in Pre-training Data Makes Better Diffusion ModelsHao Chen, Yujin Han, Diganta Misra, Xiang Li et al.NeurIPS 2024 · 14 citations
- Bayesian-guided Label Mapping for Visual ReprogrammingChengyi Cai, Zesheng Ye, Lei Feng, Jianzhong Qi et al.NeurIPS 2024 · 14 citations
- Exploring the Noise Robustness of Online Conformal PredictionHuajun Xi, Kangdao Liu, Hao Zeng, Wenguang Sun et al.NeurIPS 2025 · 4 citations
- From Pretraining to Pathology: How Noise Leads to Catastrophic Inheritance in Medical ModelsHao Sun, Zhongyi Han, Hao Chen, Jindong Wang et al.NeurIPS 2025 · 4 citations
- Mastering Collaborative Multi-Modal Data Selection: A Focus on Informativeness, Uniqueness, and RepresentativenessQifan Yu, Zhebei Shen, Zhongqi Yue, Yang Wu et al.ICCV 2025 · 1 citation
Builds on50
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
Related papers
- Improved OOD Generalization via Adversarial Training and PretraingMingyang Yi, Lu Hou, Jiacheng Sun, Lifeng Shang et al.ICML 2021 · 99 citations
- Rethinking Supervised Pre-Training for Better Downstream TransferringYutong Feng, Jianwen Jiang, Mingqian Tang, Rong Jin et al.ICLR 2022 · 51 citations
- FairTune: Optimizing Parameter Efficient Fine Tuning for Fairness in Medical Image AnalysisRaman Dutt, Ondrej Bohdal, Sotirios A. Tsaftaris, Timothy M. HospedalesICLR 2024 · 30 citations
- On the Connection between Pre-training Data Diversity and Fine-tuning RobustnessVivek Ramanujan, Thao Nguyen, Sewoong Oh, Ali Farhadi et al.NeurIPS 2023 · 40 citations
- Detecting Label Errors by Using Pre-Trained Language ModelsDerek Chong, Jenny Hong, Christopher D. ManningEMNLP 2022 · 8 citations
