LoRA Training in the NTK Regime has No Spurious Local Minima
Uijeong Jang, Jason D. Lee, Ernest K. Ryu
2024年份
41被引次数
21顶会引用
摘要
Low-rank adaptation (LoRA) has become the standard approach for parameter-efficient fine-tuning of large language models (LLM), but our theoretical understanding of LoRA has been limited. In this work, we theoretically analyze LoRA fine-tuning in the neural tangent kernel (NTK) regime with data points, showing: (i) full fine-tuning (without LoRA) admits a low-rank solution of rank ; (ii) using LoRA with rank eliminates spurious local minima, allowing gradient descent to find the low-rank solutions; (iii) the low-rank solution found using LoRA generalizes well.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Understanding Linear Probing then Fine-tuning Language Models from NTK PerspectiveAkiyoshi Tomihari, Issei SatoNeurIPS 2024 · 被引用 30 次
- Merge before Forget: A Single LoRA Continual Learning via Continual MergingFuli Qiao, Mehrdad MahdaviICLR 2026 · 被引用 11 次
- On the Convergence Rate of LoRA Gradient DescentSiqiao Mu, Diego KlabjanICML 2026 · 被引用 8 次
- The Primacy of Magnitude in Low-Rank AdaptationZicheng Zhang, Haoran Li, Yifeng Zhang, Guoqiang Gong 等NeurIPS 2025 · 被引用 7 次
- Linearization Explains Fine-Tuning in Large Language ModelsZahra Rahimi Afzal, Tara Esmaeilbeig, Mojtaba Soltanalian, Mesrob I. OhannessianNeurIPS 2025 · 被引用 5 次
它引用的顶会 Paper21
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi 等ICLR 2020 · 被引用 481 次
- Why are Adaptive Methods Good for Attention Models?Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim 等NeurIPS 2020 · 被引用 397 次
相关 Paper
- The Expressive Power of Low-Rank AdaptationYuchen Zeng, Kangwook LeeICLR 2024 · 被引用 116 次
- LoRA Training Provably Converges to a Low-Rank Global Minimum Or It Fails Loudly (But it Probably Won't Fail)Junsu Kim, Jaeyeon Kim, Ernest K. RyuICML 2025
- Does Low Rank Adaptation Lead to Lower Robustness against Training-Time Attacks?Zi Liang, Haibo Hu, Qingqing Ye, Yaxin Xiao 等ICML 2025
- LoRA-Pro: Are Low-Rank Adapters Properly Optimized?Zhengbo Wang, Jian Liang, Ran He, Zilei Wang 等ICLR 2025
- MELoRA: Mini-Ensemble Low-Rank Adapters for Parameter-Efficient Fine-TuningPengjie Ren, Chengshun Shi, Shiguang Wu, Mengqi Zhang 等ACL 2024
