Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks
Hao Chen, Jindong Wang, Ankit Shah, Ran Tao, Hongxin Wei, Xing Xie, Masashi Sugiyama, Bhiksha Raj
摘要
Pre-training on large-scale datasets and then fine-tuning on downstream tasks have become a standard practice in deep learning. However, pre-training datasets, while inaccessible or too expensive to handle, often contain label noise that may adversely affect the generalization of the model and pose unexpected risks. This paper aims to understand the nature of noise in pre-training datasets and then mitigate its impact on downstream tasks. Specifically, through extensive experiments of supervised pre-training models on synthetic noisy ImageNet-1K and YFCC15M datasets, we demonstrate that while slight noise in pre-training can benefit in-domain (ID) performance, where the training and testing data share the same distribution, it always deteriorates out-of-domain (OOD) performance, where training and testing distributions are different. We empirically ascertain that the reason behind is noise in pre-training shapes the feature space differently. We then propose a light-weight black-box tuning method (NMTune) to affine the feature space to mitigate the malignant effect of noise and improve generalization on both ID and OOD tasks, considering that one may not be able to access or fully fine-tune the pre-trained models. We conduct extensive experiments on popular vision and language models including APIs that are supervised and self-supervised pre-trained on real data for evaluation. Our results show the importance of this novel and fundamental research direction, which we term Noisy Model Learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Slight Corruption in Pre-training Data Makes Better Diffusion ModelsHao Chen, Yujin Han, Diganta Misra, Xiang Li 等NeurIPS 2024 · 被引用 14 次
- Bayesian-guided Label Mapping for Visual ReprogrammingChengyi Cai, Zesheng Ye, Lei Feng, Jianzhong Qi 等NeurIPS 2024 · 被引用 14 次
- Exploring the Noise Robustness of Online Conformal PredictionHuajun Xi, Kangdao Liu, Hao Zeng, Wenguang Sun 等NeurIPS 2025 · 被引用 4 次
- From Pretraining to Pathology: How Noise Leads to Catastrophic Inheritance in Medical ModelsHao Sun, Zhongyi Han, Hao Chen, Jindong Wang 等NeurIPS 2025 · 被引用 4 次
- Mastering Collaborative Multi-Modal Data Selection: A Focus on Informativeness, Uniqueness, and RepresentativenessQifan Yu, Zhebei Shen, Zhongqi Yue, Yang Wu 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper50
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
相关 Paper
- Improved OOD Generalization via Adversarial Training and PretraingMingyang Yi, Lu Hou, Jiacheng Sun, Lifeng Shang 等ICML 2021 · 被引用 99 次
- Rethinking Supervised Pre-Training for Better Downstream TransferringYutong Feng, Jianwen Jiang, Mingqian Tang, Rong Jin 等ICLR 2022 · 被引用 51 次
- FairTune: Optimizing Parameter Efficient Fine Tuning for Fairness in Medical Image AnalysisRaman Dutt, Ondrej Bohdal, Sotirios A. Tsaftaris, Timothy M. HospedalesICLR 2024 · 被引用 30 次
- On the Connection between Pre-training Data Diversity and Fine-tuning RobustnessVivek Ramanujan, Thao Nguyen, Sewoong Oh, Ali Farhadi 等NeurIPS 2023 · 被引用 40 次
- Detecting Label Errors by Using Pre-Trained Language ModelsDerek Chong, Jenny Hong, Christopher D. ManningEMNLP 2022 · 被引用 8 次
