Sparse is Enough in Fine-tuning Pre-trained Large Language Models
Weixi Song, Zuchao Li, Lefei Zhang, Hai Zhao, Bo Du
摘要
With the prevalence of pre-training-fine-tuning paradigm, how to efficiently adapt the pre-trained model to the downstream tasks has been an intriguing issue. Parameter-Efficient Fine-Tuning (PEFT) methods have been proposed for low-cost adaptation. Although PEFT has demonstrated effectiveness and been widely applied, the underlying principles are still unclear. In this paper, we adopt the PAC-Bayesian generalization error bound, viewing pre-training as a shift of prior distribution which leads to a tighter bound for generalization error. We validate this shift from the perspectives of oscillations in the loss landscape and the quasi-sparsity in gradient distribution. Based on this, we propose a gradient-based sparse fine-tuning algorithm, named Sparse Increment Fine-Tuning (SIFT), and validate its effectiveness on a range of tasks including the GLUE Benchmark and Instruction-tuning. The code is accessible at https://github.com/song-wx/SIFT/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Towards Understanding the Dynamics of Low-Rank AdaptationShu Ding, Yang Peng, Hangan Zhou, Xinyu Lu 等ICML 2026 · 被引用 13 次
- How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise GradientsMing Li, Yanhong Li, Ziyue Li, Tianyi ZhouACL 2026 · 被引用 9 次
- MGUP: A Momentum-Gradient Alignment Update Policy for Stochastic OptimizationDa Chang, Ganzhao YuanNeurIPS 2025 · 被引用 9 次
- Efficient Orthogonal Fine-Tuning with Principal Subspace AdaptationFei Wu, Jia Hu, Geyong Min, Shiqiang WangICLR 2026 · 被引用 5 次
- Learn from Downstream and Be Yourself in Multimodal Large Language Models Fine-TuningWenke Huang, Jian Liang, Zekun Shi, Didi Zhu 等ICML 2025 · 被引用 1 次
它引用的顶会 Paper6
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
- Full Parameter Fine-tuning for Large Language Models with Limited ResourcesKai Lv, Yuqing Yang, Tengxiao Liu, Qipeng Guo 等ACL 2024 · 被引用 61 次
- Adaptive Budget Allocation for Parameter-Efficient Fine-TuningQingru Zhang, Minshuo Chen, Alexander Bukharin, Pengcheng He 等ICLR 2023 · 被引用 32 次
相关 Paper
- GEM: A Scale-Aware and Distribution-Sensitive Sparse Fine-Tuning Framework for Effective Downstream AdaptationSungmin Kang, Jisoo Kim, Salman Avestimehr, Sunwoo LeeAAAI 2026
- PAC-tuning: Fine-tuning Pre-trained Language Models with PAC-driven Perturbed Gradient DescentGuangliang Liu, Zhiyu Xue, Xitong Zhang, Kristen Marie Johnson 等EMNLP 2023 · 被引用 1 次
- Parameter-Efficient Fine-Tuning without Introducing New LatencyBaohao Liao, Yan Meng, Christof MonzACL 2023 · 被引用 26 次
- Flat-LoRA: Low-Rank Adaptation over a Flat Loss LandscapeTao Li, Zhengbao He, Yujun Li, Yasheng Wang 等ICML 2025
- RoseLoRA: Row and Column-wise Sparse Low-rank Adaptation of Pre-trained Language Model for Knowledge Editing and Fine-tuningHaoyu Wang, Tianci Liu, Ruirui Li, Monica Xiao Cheng 等EMNLP 2024 · 被引用 6 次
