Robust Fine-Tuning of Deep Neural Networks with Hessian-based Generalization Guarantees
Haotian Ju, Dongyue Li, Hongyang R. Zhang
摘要
We consider fine-tuning a pretrained deep neural network on a target task. We study the generalization properties of fine-tuning to understand the problem of overfitting, which has often been observed (e.g., when the target dataset is small or when the training labels are noisy). Existing generalization measures for deep networks depend on notions such as distance from the initialization (i.e., the pretrained network) of the fine-tuned model and noise stability properties of deep networks. This paper identifies a Hessian-based distance measure through PAC-Bayesian analysis, which is shown to correlate well with observed generalization gaps of fine-tuned models. Theoretically, we prove Hessian distance-based generalization bounds for fine-tuned models. We also describe an extended study of fine-tuning against label noise, where overfitting remains a critical problem. We present an algorithm and a generalization error guarantee for this algorithm under a class conditional independent noise model. Empirically, we observe that the Hessian-based distance measure can match the scale of the observed generalization gap of fine-tuned models in practice. We also test our algorithm on several image classification tasks with noisy training labels, showing gains over prior methods and decreases in the Hessian distance measure of the fine-tuned model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Improved Regularization and Robustness for Fine-tuning in Neural NetworksDongyue Li, Hongyang R. ZhangNeurIPS 2021 · 被引用 76 次
- Deep learning with kernels through RKHM and the Perron-Frobenius operatorYuka Hashimoto, Masahiro Ikeda, Hachem KadriNeurIPS 2023 · 被引用 13 次
- Pre-training LLM without Learning Rate Decay Enhances Supervised Fine-TuningKazuki Yano, Shun Kiyono, Sosuke Kobayashi, Sho Takase 等ICLR 2026 · 被引用 13 次
- A Theoretical Analysis of the Test Error of Finite-Rank Kernel Ridge RegressionTin Sum Cheng, Aurélien Lucchi, Anastasis Kratsios, Ivan Dokmanic 等NeurIPS 2023 · 被引用 11 次
- Efficient Ensemble for Fine-tuning Language Models on Multiple DatasetsDongyue Li, Ziniu Zhang, Lu Wang, Hongyang R. ZhangACL 2025 · 被引用 8 次
它引用的顶会 Paper21
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Early-Learning Regularization Prevents Memorization of Noisy LabelsSheng Liu, Jonathan Niles-Weed, Narges Razavian, Carlos Fernandez-GrandaNeurIPS 2020 · 被引用 798 次
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- What is being transferred in transfer learning?Behnam Neyshabur, Hanie Sedghi, Chiyuan ZhangNeurIPS 2020 · 被引用 654 次
相关 Paper
- Simple and Effective Regularization Methods for Training on Noisily Labeled Data with Generalization GuaranteeWei Hu, Zhiyuan Li, Dingli YuICLR 2020 · 被引用 140 次
- RATT: Leveraging Unlabeled Data to Guarantee GeneralizationSaurabh Garg, Sivaraman Balakrishnan, J. Zico Kolter, Zachary C. LiptonICML 2021 · 被引用 30 次
- Normalized Flat Minima: Exploring Scale Invariant Definition of Flat Minima for Neural Networks Using PAC-Bayesian AnalysisYusuke Tsuzuku, Issei Sato, Masashi SugiyamaICML 2020 · 被引用 91 次
- Improved Fine-Tuning by Better Leveraging Pre-Training DataZiquan Liu, Yi Xu, Yuanhong Xu, Qi Qian 等NeurIPS 2022 · 被引用 69 次
- Distance-Based Regularisation of Deep Networks for Fine-TuningHenry Gouk, Timothy M. Hospedales, Massimiliano PontilICLR 2021 · 被引用 65 次
