Measuring the Instability of Fine-Tuning
Yupei Du, Dong Nguyen
Abstract
Fine-tuning pre-trained language models on downstream tasks with varying random seeds has been shown to be unstable, especially on small datasets. Many previous studies have investigated this instability and proposed methods to mitigate it. However, most studies only used the standard deviation of performance scores (SD) as their measure, which is a narrow characterization of instability. In this paper, we analyze SD and six other measures quantifying instability at different levels of granularity. Moreover, we propose a systematic framework to evaluate the validity of these measures. Finally, we analyze the consistency and difference between different measures by reassessing existing instability mitigation methods. We hope our results will inform the development of better measurements of fine-tuning instability. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- OraPO: Oracle-educated Reinforcement Learning for Data-efficient and Factual Radiology Report GenerationZhuoxiao Chen, Hongyang Yu, Ying Xu, Yadan Luo et al.CVPR 2026 · 3 citations
- Mixtures of In-Context LearnersGiwon Hong, Emile van Krieken, Edoardo Maria Ponti, Nikolay Malkin et al.ACL 2025
- PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training RunsOskar van der Wal, Pietro Lesci, Max Müller-Eberstein, Naomi Saphra et al.ICLR 2025
Builds on9
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- On the Stability of Fine-tuning BERT: Misconceptions, Explanations, and Strong BaselinesMarius Mosbach, Maksym Andriushchenko, Dietrich KlakowICLR 2021 · 448 citations
- Mixout: Effective Regularization to Finetune Large-scale Pretrained Language ModelsCheolhyoung Lee, Kyunghyun Cho, Wanmo KangICLR 2020 · 233 citations
- On Linear Identifiability of Learned RepresentationsGeoffrey Roeder, Luke Metz, Durk KingmaICML 2021 · 107 citations
Related papers
- Same Pre-training Loss, Better Downstream: Implicit Bias Matters for Language ModelsHong Liu, Sang Michael Xie, Zhiyuan Li, Tengyu MaICML 2023 · 82 citations
- PTP: Boosting Stability and Performance of Prompt Tuning with Perturbation-Based RegularizerLichang Chen, Jiuhai Chen, Heng Huang, Minhao ChengEMNLP 2023 · 3 citations
- Fine-tuning can Help Detect Pretraining Data from Large Language ModelsHengxiang Zhang, Songxin Zhang, Bingyi Jing, Hongxin WeiICLR 2025
- ROSE: Robust Selective Fine-tuning for Pre-trained Language ModelsLan Jiang, Hao Zhou, Yankai Lin, Peng Li et al.EMNLP 2022 · 5 citations
- Revisiting Parameter-Efficient Tuning: Are We Really There Yet?Guanzheng Chen, Fangyu Liu, Zaiqiao Meng, Shangsong LiangEMNLP 2022 · 49 citations
