A Closer Look at How Fine-tuning Changes BERT
Yichu Zhou, Vivek Srikumar
摘要
Given the prevalence of pre-trained contextualized representations in today’s NLP, there have been many efforts to understand what information they contain, and why they seem to be universally successful. The most common approach to use these representations involves fine-tuning them for an end task. Yet, how fine-tuning changes the underlying embedding space is less studied. In this work, we study the English BERT family and use two probing techniques to analyze how fine-tuning changes the space. We hypothesize that fine-tuning affects classification performance by increasing the distances between examples associated with different labels. We confirm this hypothesis with carefully designed experiments on five different NLP tasks. Via these experiments, we also discover an exception to the prevailing wisdom that “fine-tuning always improves performance”. Finally, by comparing the representations before and after fine-tuning, we discover that fine-tuning does not introduce arbitrary changes to representations; instead, it adjusts the representations to downstream tasks while largely preserving the original spatial structure of the data points.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity TrackingNikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov 等ICLR 2024 · 被引用 113 次
- Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-TuningJulian Minder, Clément Dumas, Caden Juang, Bilal Chughtai 等NeurIPS 2025 · 被引用 32 次
- CogTaskonomy: Cognitively Inspired Task Taxonomy Is Beneficial to Transfer Learning in NLPYifei Luo, Minghui Xu, Deyi XiongACL 2022 · 被引用 20 次
- Vision-Language Model Fine-Tuning via Simple Parameter-Efficient ModificationMing Li, Jike Zhong, Chenxin Li, Liuzhuozheng Li 等EMNLP 2024 · 被引用 18 次
- Superficial Self-Improved Reasoners Benefit from Model MergingXiangchi Yuan, Chunhui Zhang, Zheyuan Liu, Dachuan Shi 等EMNLP 2025 · 被引用 15 次
它引用的顶会 Paper10
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- On the Stability of Fine-tuning BERT: Misconceptions, Explanations, and Strong BaselinesMarius Mosbach, Maksym Andriushchenko, Dietrich KlakowICLR 2021 · 被引用 448 次
- Revisiting Few-sample BERT Fine-tuningTianyi Zhang, Felix Wu, Arzoo Katiyar, Kilian Q. Weinberger 等ICLR 2021 · 被引用 172 次
- Intermediate-Task Transfer Learning with Pretrained Language Models: When and Why Does It Work?Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut 等ACL 2020 · 被引用 168 次
- Perturbed Masking: Parameter-free Probing for Analyzing and Interpreting BERTZhiyong Wu, Yun Chen, Ben Kao, Qun LiuACL 2020 · 被引用 158 次
相关 Paper
- On the Transformation of Latent Space in Fine-Tuned NLP ModelsNadir Durrani, Hassan Sajjad, Fahim Dalvi, Firoj AlamEMNLP 2022 · 被引用 3 次
- Spying on Your Neighbors: Fine-grained Probing of Contextual Embeddings for Information about Surrounding WordsJosef Klafka, Allyson EttingerACL 2020 · 被引用 2 次
- Probing BERT in Hyperbolic SpacesBoli Chen, Yao Fu, Guangwei Xu, Pengjun Xie 等ICLR 2021 · 被引用 19 次
- Exploring the Role of BERT Token Representations to Explain Sentence Probing ResultsHosein Mohebbi, Ali Modarressi, Mohammad Taher PilehvarEMNLP 2021 · 被引用 14 次
- Fine-Tuning is Fine, if CalibratedZheda Mai, Arpita Chowdhury, Ping Zhang, Cheng-Hao Tu 等NeurIPS 2024 · 被引用 34 次
