Rethinking the Hyperparameters for Fine-tuning
Hao Li, Pratik Chaudhari, Hao Yang, Michael Lam, Avinash Ravichandran, Rahul Bhotika, Stefano Soatto
摘要
Fine-tuning from pre-trained ImageNet models has become the de-facto standard for various computer vision tasks. Current practices for fine-tuning typically involve selecting an ad-hoc choice of hyper-parameters and keeping them fixed to values normally used for training from scratch. This paper re-examines several common practices of setting hyper-parameters for fine-tuning. Our findings are based on extensive empirical evaluation for fine-tuning on various transfer learning benchmarks. (1) While prior works have thoroughly investigated learning rate and batch size, momentum for fine-tuning is a relatively unexplored parameter. We find that picking the right value for momentum is critical for fine-tuning performance and connect it with previous theoretical findings. (2) Optimal hyper-parameters for fine-tuning in particular the effective learning rate are not only dataset dependent but also sensitive to the similarity between the source domain and target domain. This is in contrast to hyper-parameters for training from scratch. (3) Reference-based regularization that keeps models close to the initial model does not necessarily apply for "dissimilar" datasets. Our findings challenge common practices of fine- tuning and encourages deep learning practitioners to rethink the hyper-parameters for fine-tuning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper36
- Robust fine-tuning of zero-shot modelsMitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li 等CVPR 2022 · 被引用 364 次
- LogME: Practical Assessment of Pre-trained Models for Transfer LearningKaichao You, Yong Liu, Jianmin Wang, Mingsheng LongICML 2021 · 被引用 253 次
- FREE: Feature Refinement for Generalized Zero-Shot LearningShiming Chen, Wenjie Wang, Beihao Xia, Qinmu Peng 等ICCV 2021 · 被引用 171 次
- Pretraining Representations for Data-Efficient Reinforcement LearningMax Schwarzer, Nitarshan Rajkumar, Michael Noukhovitch, Ankesh Anand 等NeurIPS 2021 · 被引用 151 次
- Flora: Low-Rank Adapters Are Secretly Gradient CompressorsYongchang Hao, Yanshuai Cao, Lili MouICML 2024 · 被引用 113 次
它引用的顶会 Paper2
相关 Paper
- Demystify Hyperparameters for Stochastic Optimization with Transferable RepresentationsJianhui Sun, Mengdi Huai, Kishlay Jha, Aidong ZhangKDD 2022 · 被引用 5 次
- Improved Fine-Tuning by Better Leveraging Pre-Training DataZiquan Liu, Yi Xu, Yuanhong Xu, Qi Qian 等NeurIPS 2022 · 被引用 69 次
- Co-Tuning for Transfer LearningKaichao You, Zhi Kou, Mingsheng Long, Jianmin WangNeurIPS 2020 · 被引用 105 次
- When, Why, and Which Pretrained GANs Are Useful?Timofey Grigoryev, Andrey Voynov, Artem BabenkoICLR 2022 · 被引用 25 次
- What Makes Transfer Learning Work for Medical Images: Feature Reuse & Other FactorsChristos Matsoukas, Johan Fredin Haslum, Moein Sorkhei, Magnus Söderberg 等CVPR 2022 · 被引用 93 次
