Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning
Armen Aghajanyan, Sonal Gupta, Luke Zettlemoyer
Abstract
Although pretrained language models can be fine-tuned to produce state-of-theart results for a very wide range of language understanding tasks, the dynamics of this process are not well understood, especially in the low data regime. Why can we use relatively vanilla gradient descent algorithms (e.g., without strong regularization) to tune a model with hundreds of millions of parameters on datasets with only hundreds or thousands of labeled examples? In this paper, we argue that analyzing fine-tuning through the lens of intrinsic dimension provides us with empirical and theoretical intuitions to explain this remarkable phenomenon. We empirically show that common pre-trained models have a very low intrinsic dimension; in other words, there exists a low dimension reparameterization that is as effective for fine-tuning as the full parameter space. For example, by optimizing only 200 trainable parameters randomly projected back into the full space, we can tune a RoBERTa model to achieve 90% of the full parameter performance levels on MRPC. Furthermore, we empirically show that pre-training implicitly minimizes intrinsic dimension and, perhaps surprisingly, larger models tend to have lower intrinsic dimension after a fixed number of pre-training updates, at least in part explaining their extreme effectiveness. Lastly, we connect intrinsic dimensionality with low dimensional task representations and compression based generalization bounds to provide intrinsic-dimension-based generalization bounds that are independent of the full parameter count.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers271
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context LearningHaokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta et al.NeurIPS 2022 · 1,483 citations
- Compacter: Efficient Low-Rank Hypercomplex Adapter LayersRabeeh Karimi Mahabadi, James Henderson, Sebastian RuderNeurIPS 2021 · 700 citations
- Fine-Tuning Language Models with Just Forward PassesSadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian et al.NeurIPS 2023 · 495 citations
Builds on5
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- The Lottery Ticket Hypothesis for Pre-trained BERT NetworksTianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu et al.NeurIPS 2020 · 428 citations
- When BERT Plays the Lottery, All Tickets Are WinningSai Prasanna, Anna Rogers, Anna RumshiskyEMNLP 2020 · 114 citations
- Better Fine-Tuning by Reducing Representational CollapseArmen Aghajanyan, Akshat Shrivastava, Anchit Gupta, Naman Goyal et al.ICLR 2021 · 20 citations
Related papers
- Fine-tuning Happens in Tiny Subspaces: Exploring Intrinsic Task-specific Subspaces of Pre-trained Language ModelsZhong Zhang, Bang Liu, Junming ShaoACL 2023 · 8 citations
- IsoBN: Fine-Tuning BERT with Isotropic Batch NormalizationWenxuan Zhou, Bill Yuchen Lin, Xiang RenAAAI 2021 · 29 citations
- Gradient Intrinsic Dimensionality Alignment:Narrowing The Gap Between Low-Rank Adaptation and Full Fine-TuningJingqi Ye, Haonan He, Minglei Li, Fujun Han et al.ICLR 2026
- Bridging Information-Theoretic and Geometric Compression in Language ModelsEmily Cheng, Corentin Kervadec, Marco BaroniEMNLP 2023 · 5 citations
- Non-Vacuous Generalization Bounds for Large Language ModelsSanae Lotfi, Marc Anton Finzi, Yilun Kuang, Tim G. J. Rudner et al.ICML 2024 · 49 citations
