Low-cost Full Fine-tuning: Learning What to Update for LLMs
Li, Yaming Guo, Shenghao Gao, Xinlong Chen, Zuhao Xu, Ying Sun, Chao Wang, Hui Xiong
摘要
While Large language models (LLMs) have strong abilities, they generally rely on fine-tuning to supplement downstream task-specific knowledge. Due to the prohibitive memory overhead of full fine-tuning (FT), existing parameter-efficient fine-tuning techniques, e.g., LoRA and Adapters, update parameters only in low-rank or restricted subspaces. However, they fail to approximate FT---the performative fine-tuner---and risk performance degradation in tough tasks. Therefore, we naturally raise a Low-cost Full Fine-tuning question: Can we approach standard full fine-tuning in theory, yet with much lower costs in practice? Our key insight is that performing selective updates at each step can, theoretically, recover FT asymptotically, while being cost-effective and ignoring no parameter direction. This motivates a new general fine-tuning paradigm (called Think-Touch ): we first predict potentials of parameter groups ( think ) and then update only the selected ( touch ) in one step. Theoretically, we show that under a very weak sufficient condition---divergence of the cumulative coverage of the expected gradient norm---any selection strategy can converge in the full-parameter space to a stationary point at which the FT admits no further first-order improvement. Besides, we further derive the general convergence rate for our paradigm and identify a post-hoc greedy strategy that is rate-optimal. Unfortunately, this strategy cannot be directly applied in practice due to its reliance on full and accurate gradient information. Thus, we propose a bandit-based method to online approximate this ideal strategy in the long run with a rigorous regret guarantee. Extensive experimental results on various tasks demonstrate the potential of our paradigm, including much lower space overheads against FT and better performance than LoRAs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context LearningHaokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta 等NeurIPS 2022 · 被引用 1,483 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- Compacter: Efficient Low-Rank Hypercomplex Adapter LayersRabeeh Karimi Mahabadi, James Henderson, Sebastian RuderNeurIPS 2021 · 被引用 700 次
相关 Paper
- LoRA Training in the NTK Regime has No Spurious Local MinimaUijeong Jang, Jason D. Lee, Ernest K. RyuICML 2024 · 被引用 41 次
- Parameter-efficient Tuning for Large Language Model without Calculating Its GradientsFeihu Jin, Jiajun Zhang, Chengqing ZongEMNLP 2023 · 被引用 2 次
- MELoRA: Mini-Ensemble Low-Rank Adapters for Parameter-Efficient Fine-TuningPengjie Ren, Chengshun Shi, Shiguang Wu, Mengqi Zhang 等ACL 2024
- HiRA: Parameter-Efficient Hadamard High-Rank Adaptation for Large Language ModelsQiushi Huang, Tom Ko, Zhan Zhuang, Lilian Tang 等ICLR 2025
- ScaLoRA: Optimally Scaled Low-Rank Adaptation for Efficient High-Rank Fine-TuningYilang Zhang, Xiaodong Yang, Yiwei Cai, Georgios B. GiannakisICML 2026 · 被引用 1 次
