LoRA-One: One-Step Full Gradient Could Suffice for Fine-Tuning Large Language Models, Provably and Efficiently
Yuanhe Zhang, Fanghui Liu, Yudong Chen
摘要
This paper explores how theory can guide and enhance practical algorithms, using Low-Rank Adaptation (LoRA) (Hu et al., 2022) in large language models as a case study. We rigorously prove that, under gradient descent, LoRA adapters align with specific singular subspaces of the onestep full fine-tuning gradient. This result suggests that, by properly initializing the adapters using the one-step full gradient, subspace alignment can be achieved immediately-applicable to both linear and nonlinear models. Building on our theory, we propose a theory-driven algorithm, LoRA-One, where the linear convergence (as well as generalization) is built and incorporating preconditioners theoretically helps mitigate the effects of ill-conditioning. Besides, our theory reveals connections between LoRA-One and other gradient-alignment-based methods, helping to clarify misconceptions in the design of such algorithms. LoRA-One achieves significant empirical improvements over LoRA and its variants across benchmarks in natural language understanding, mathematical reasoning, and code generation. Code is available at: https://github. com/YuanheZ/LoRA-One .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Towards Understanding the Dynamics of Low-Rank AdaptationShu Ding, Yang Peng, Hangan Zhou, Xinyu Lu 等ICML 2026 · 被引用 13 次
- On the Convergence Rate of LoRA Gradient DescentSiqiao Mu, Diego KlabjanICML 2026 · 被引用 8 次
- ABBA-Adapters: Efficient and Expressive Fine-Tuning of Foundation ModelsRaghav Singhal, Kaustubh Ponkshe, Rohit Vartak, Praneeth VepakommaICLR 2026 · 被引用 3 次
- Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable DataHancheng Min, Zhihui Zhu, René VidalNeurIPS 2025 · 被引用 3 次
- DiaBlo: Diagonal Blocks Are Sufficient For FinetuningSelcuk Gurses, Aozhong Zhang, Yanxia Deng, Xun Dong 等ICLR 2026 · 被引用 2 次
它引用的顶会 Paper24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- LoRA+: Efficient Low Rank Adaptation of Large ModelsSoufiane Hayou, Nikhil Ghosh, Bin YuICML 2024 · 被引用 388 次
相关 Paper
- AdaRankGrad: Adaptive Gradient Rank and Moments for Memory-Efficient LLMs Training and Fine-TuningYehonathan Refael, Jonathan Svirsky, Boris Shustin, Wasim Huleihel 等ICLR 2025
- LoRA-Pro: Are Low-Rank Adapters Properly Optimized?Zhengbo Wang, Jian Liang, Ran He, Zilei Wang 等ICLR 2025
- When Gradient Boosting Meets Adapter: Exploring Weak Learners for Parameter-Efficient Fine-tuning of LLMsYifei Zhang, Hao Zhu, Haoran Shi, Junhao Dong 等KDD 2026
- LoRA Training in the NTK Regime has No Spurious Local MinimaUijeong Jang, Jason D. Lee, Ernest K. RyuICML 2024 · 被引用 41 次
- Low Kruskal-Rank AdaptationYixing Xu, Guanchen Li, Chao Li, Xuanwu Yin 等ICML 2026
