IsoBN: Fine-Tuning BERT with Isotropic Batch Normalization
Wenxuan Zhou, Bill Yuchen Lin, Xiang Ren
Abstract
Fine-tuning pre-trained language models (PTLMs), such as BERT and its better variant RoBERTa, has been a common practice for advancing performance in natural language understanding (NLU) tasks. Recent advance in representation learning shows that isotropic (i.e., unit-variance and uncorrelated) embeddings can significantly improve performance on downstream tasks with faster convergence and better generalization. The isotropy of the pre-trained embeddings in PTLMs, however, is relatively under-explored. In this paper, we analyze the isotropy of the pre-trained [CLS] embeddings of PTLMs with straightforward visualization, and point out two major issues: high variance in their standard deviation, and high correlation between different dimensions. We also propose a new network regularization method, isotropic batch normalization (IsoBN) to address the issues, towards learning more isotropic representations in fine-tuning by dynamically penalizing dominating principal components. This simple yet effective fine-tuning method yields about 1.0 absolute increment on the average of seven NLU tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b51ebf2b-af65-47f6-8e3c-dadaab8258e7Cited by top-tier papers3
- Stable Anisotropic RegularizationWilliam Rudman, Carsten EickhoffICLR 2024 · 13 citations
- Differential Privacy, Linguistic Fairness, and Training Data Influence: Impossibility and Possibility Theorems for Multilingual Language ModelsPhillip Rust, Anders SøgaardICML 2023 · 7 citations
- Reliable Measures of Spread in High Dimensional Latent SpacesAnna C. Marbut, Katy McKinney-Bock, Travis J. WheelerICML 2023 · 5 citations
Builds on3
- FreeLB: Enhanced Adversarial Training for Natural Language UnderstandingChen Zhu, Yu Cheng, Zhe Gan, Siqi Sun et al.ICLR 2020 · 502 citations
- SMART: Robust and Efficient Fine-Tuning for Pre-trained Natural Language Models through Principled Regularized OptimizationHaoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu et al.ACL 2020 · 148 citations
- Improving Neural Language Generation with Spectrum ControlLingxiao Wang, Jing Huang, Kevin Huang, Ziniu Hu et al.ICLR 2020 · 94 citations
Related papers
- Positional Artefacts Propagate Through Masked Language Model EmbeddingsZiyang Luo, Artur Kulmizev, Xiaoxi MaoACL 2021
- Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-TuningArmen Aghajanyan, Sonal Gupta, Luke ZettlemoyerACL 2021
- Mixout: Effective Regularization to Finetune Large-scale Pretrained Language ModelsCheolhyoung Lee, Kyunghyun Cho, Wanmo KangICLR 2020 · 233 citations
- Kernel-Whitening: Overcome Dataset Bias with Isotropic Sentence EmbeddingSongyang Gao, Shihan Dou, Qi Zhang, Xuanjing HuangEMNLP 2022 · 10 citations
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang et al.EMNLP 2020 · 538 citations
