Improved Regularization and Robustness for Fine-tuning in Neural Networks
Dongyue Li, Hongyang R. Zhang
摘要
A widely used algorithm for transfer learning is fine-tuning, where a pre-trained model is fine-tuned on a target task with a small amount of labeled data. When the capacity of the pre-trained model is significantly larger than the size of the target dataset, fine-tuning is prone to overfitting and memorizing the training labels. Hence, a crucial question is to regularize fine-tuning and ensure its robustness against noise. To address this question, we begin by analyzing the generalization properties of fine-tuning. We present a PAC-Bayes generalization bound that depends on the distance traveled in each layer during fine-tuning and the noise stability of the fine-tuned model. We empirically measure these quantities. Based on the analysis, we propose regularized self-labeling -- the interpolation between regularization and self-labeling methods, including (i) layer-wise regularization to constrain the distance traveled in each layer; (ii) self-label-correction and label-reweighting to correct mislabeled data points (that the model is confident) and reweight less confident data points. We validate our approach on an extensive collection of image and text datasets using multiple pre-trained model architectures. Our approach improves baseline methods by 1.76% (on average) for seven image classification tasks and 0.75% for a few-shot classification task. When the target data set includes noisy labels, our approach outperforms baseline methods by an average of 3.56% in two noisy settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- On the Effectiveness of Parameter-Efficient Fine-TuningZihao Fu, Haoran Yang, Anthony Man-Cho So, Wai Lam 等AAAI 2023 · 被引用 234 次
- Robust Fine-Tuning of Deep Neural Networks with Hessian-based Generalization GuaranteesHaotian Ju, Dongyue Li, Hongyang R. ZhangICML 2022 · 被引用 41 次
- CLAP4CLIP: Continual Learning with Probabilistic Finetuning for Vision-Language ModelsSaurav Jha, Dong Gong, Lina YaoNeurIPS 2024 · 被引用 36 次
- Fine-Tuning is Fine, if CalibratedZheda Mai, Arpita Chowdhury, Ping Zhang, Cheng-Hao Tu 等NeurIPS 2024 · 被引用 34 次
- PACE: Marrying generalization in PArameter-efficient fine-tuning with Consistency rEgularizationYao Ni, Shan Zhang, Piotr KoniuszNeurIPS 2024 · 被引用 25 次
它引用的顶会 Paper13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Symmetric Cross Entropy for Robust Learning With Noisy LabelsYisen Wang, Xingjun Ma, Zaiyi Chen, Yuan Luo 等ICCV 2019 · 被引用 1,125 次
- Early-Learning Regularization Prevents Memorization of Noisy LabelsSheng Liu, Jonathan Niles-Weed, Narges Razavian, Carlos Fernandez-GrandaNeurIPS 2020 · 被引用 798 次
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- A Baseline for Few-Shot Image ClassificationGuneet Singh Dhillon, Pratik Chaudhari, Avinash Ravichandran, Stefano SoattoICLR 2020 · 被引用 640 次
相关 Paper
- PAC-tuning: Fine-tuning Pre-trained Language Models with PAC-driven Perturbed Gradient DescentGuangliang Liu, Zhiyu Xue, Xitong Zhang, Kristen Marie Johnson 等EMNLP 2023 · 被引用 1 次
- Co-Tuning for Transfer LearningKaichao You, Zhi Kou, Mingsheng Long, Jianmin WangNeurIPS 2020 · 被引用 105 次
- Mitigating Memorization of Noisy Labels via Regularization between RepresentationsHao Cheng, Zhaowei Zhu, Xing Sun, Yang LiuICLR 2023 · 被引用 8 次
- Trainable Projected Gradient Method for Robust Fine-TuningJunjiao Tian, Xiaoliang Dai, Chih-Yao Ma, Zecheng He 等CVPR 2023
- Generalization Bounds for Meta-Learning via PAC-Bayes and Uniform StabilityAlec Farid, Anirudha MajumdarNeurIPS 2021 · 被引用 46 次
