Fine-tuning Happens in Tiny Subspaces: Exploring Intrinsic Task-specific Subspaces of Pre-trained Language Models
Zhong Zhang, Bang Liu, Junming Shao
摘要
Pre-trained language models (PLMs) are known to be overly parameterized and have significant redundancy, indicating a small degree of freedom of the PLMs. Motivated by the observation, in this paper, we study the problem of re-parameterizing and fine-tuning PLMs from a new perspective: Discovery of intrinsic task-specific subspace. Specifically, by exploiting the dynamics of the fine-tuning process for a given task, the parameter optimization trajectory is learned to uncover its intrinsic task-specific subspace. A key finding is that PLMs can be effectively fine-tuned in the subspace with a small number of free parameters. Beyond, we observe some outlier dimensions emerging during fine-tuning in the subspace. Disabling these dimensions degrades the model performance significantly. This suggests that these dimensions are crucial to induce task-specific knowledge to downstream tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-TuningJulian Minder, Clément Dumas, Caden Juang, Bilal Chughtai 等NeurIPS 2025 · 被引用 32 次
- Compressible Dynamics in Deep Overparameterized Low-Rank Learning & AdaptationCan Yaras, Peng Wang, Laura Balzano, Qing QuICML 2024 · 被引用 29 次
- Parameter Efficient Fine-tuning via Explained Variance AdaptationFabian Paischer, Lukas Hauzenberger, Thomas Schmied, Benedikt Alkin 等NeurIPS 2025 · 被引用 25 次
- Learning to Interpret Weight Differences in Language ModelsAvichal Goel, Yoon Kim, Nir N Shavit, Tony T. WangICLR 2026 · 被引用 10 次
- Uni-LoRA: One Vector is All You NeedKaiyang Li, Shaobo Han, Qing Su, Wei Li 等NeurIPS 2025 · 被引用 10 次
它引用的顶会 Paper12
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Compacter: Efficient Low-Rank Hypercomplex Adapter LayersRabeeh Karimi Mahabadi, James Henderson, Sebastian RuderNeurIPS 2021 · 被引用 700 次
- The Lottery Ticket Hypothesis for Pre-trained BERT NetworksTianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu 等NeurIPS 2020 · 被引用 428 次
- When BERT Plays the Lottery, All Tickets Are WinningSai Prasanna, Anna Rogers, Anna RumshiskyEMNLP 2020 · 被引用 114 次
- DRONE: Data-aware Low-rank Compression for Large NLP ModelsPatrick H. Chen, Hsiang-Fu Yu, Inderjit S. Dhillon, Cho-Jui HsiehNeurIPS 2021 · 被引用 109 次
相关 Paper
- Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-TuningArmen Aghajanyan, Sonal Gupta, Luke ZettlemoyerACL 2021
- Layer by Layer: Uncovering Where Multi-Task Learning Happens in Instruction-Tuned Large Language ModelsZheng Zhao, Yftah Ziser, Shay B. CohenEMNLP 2024
- A Kernel-Based View of Language Model Fine-TuningSadhika Malladi, Alexander Wettig, Dingli Yu, Danqi Chen 等ICML 2023 · 被引用 111 次
- Small Pre-trained Language Models Can be Fine-tuned as Large Models via Over-ParameterizationZe-Feng Gao, Kun Zhou, Peiyu Liu, Wayne Xin Zhao 等ACL 2023 · 被引用 3 次
- Tiny Brains, Giant Impact: Uncovering the Keystone Neurons of LLM with Just a Few PromptsXiangtian Ji, Yuxin Chen, Zhengzhou Cai, Xiang Wang 等ICML 2026
