On the Stability of Fine-tuning BERT: Misconceptions, Explanations, and Strong Baselines
Marius Mosbach, Maksym Andriushchenko, Dietrich Klakow
摘要
Fine-tuning pre-trained transformer-based language models such as BERT has become a common practice dominating leaderboards across various NLP benchmarks. Despite the strong empirical performance of fine-tuned models, fine-tuning is an unstable process: training the same model with multiple random seeds can result in a large variance of the task performance. Previous literature (Devlin et al., 2019; Lee et al., 2020; Dodge et al., 2020) identified two potential reasons for the observed instability: catastrophic forgetting and small size of the fine-tuning datasets. In this paper, we show that both hypotheses fail to explain the fine-tuning instability. We analyze BERT, RoBERTa, and ALBERT, fine-tuned on commonly used datasets from the GLUE benchmark, and show that the observed instability is caused by optimization difficulties that lead to vanishing gradients. Additionally, we show that the remaining variance of the downstream task performance can be attributed to differences in generalization where fine-tuned models with the same training loss exhibit noticeably different test performance. Based on our analysis, we present a simple but strong baseline that makes fine-tuning BERT-based models significantly more stable than the previously proposed approaches. Code to reproduce our results is available online: https://github.com/uds-lsv/bert-stable-fine-tuning .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper84
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace 等EMNLP 2020 · 被引用 1,162 次
- Supervised Contrastive Learning for Pre-trained Language Model Fine-tuningBeliz Gunel, Jingfei Du, Alexis Conneau, Veselin StoyanovICLR 2021 · 被引用 595 次
- Understanding Dataset Difficulty with V-Usable InformationKawin Ethayarajh, Yejin Choi, Swabha SwayamdiptaICML 2022 · 被引用 337 次
- On the Effectiveness of Parameter-Efficient Fine-TuningZihao Fu, Haoran Yang, Anthony Man-Cho So, Wai Lam 等AAAI 2023 · 被引用 234 次
它引用的顶会 Paper9
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen 等ICLR 2020 · 被引用 2,210 次
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng 等ICML 2020 · 被引用 1,388 次
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 被引用 1,352 次
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu 等ICLR 2020 · 被引用 1,170 次
相关 Paper
- Mixout: Effective Regularization to Finetune Large-scale Pretrained Language ModelsCheolhyoung Lee, Kyunghyun Cho, Wanmo KangICLR 2020 · 被引用 233 次
- Improved Text Classification via Contrastive Adversarial TrainingLin Pan, Chung-Wei Hang, Avirup Sil, Saloni PotdarAAAI 2022 · 被引用 115 次
- Revisiting Few-sample BERT Fine-tuningTianyi Zhang, Felix Wu, Arzoo Katiyar, Kilian Q. Weinberger 等ICLR 2021 · 被引用 172 次
- PTP: Boosting Stability and Performance of Prompt Tuning with Perturbation-Based RegularizerLichang Chen, Jiuhai Chen, Heng Huang, Minhao ChengEMNLP 2023 · 被引用 3 次
- Recall and Learn: Fine-tuning Deep Pretrained Language Models with Less ForgettingSanyuan Chen, Yutai Hou, Yiming Cui, Wanxiang Che 等EMNLP 2020 · 被引用 152 次
