LoRA Training Provably Converges to a Low-Rank Global Minimum Or It Fails Loudly (But it Probably Won't Fail)
Junsu Kim, Jaeyeon Kim, Ernest K. Ryu
Abstract
Low-rank adaptation (LoRA) has become a standard approach for fine-tuning large foundation models. However, our theoretical understanding of LoRA remains limited as prior analyses of LoRA's training dynamics either rely on linearization arguments or consider highly simplified setups. In this work, we analyze the LoRA loss landscape without such restrictive assumptions. We define two regimes: a "special regime", which includes idealized setups where linearization arguments hold, and a "generic regime" representing more realistic setups where linearization arguments do not hold. In the generic regime, we show that LoRA training converges to a global minimizer with low rank and small magnitude, or a qualitatively distinct solution with high rank and large magnitude. Finally, we argue that the zeroinitialization and weight decay in LoRA training induce an implicit bias toward the low-rank, smallmagnitude region of the parameter space-where global minima lie-thus shedding light on why LoRA training usually succeeds in finding global minima.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext befd0ce3-60a6-4eed-bab1-70cf00a24dccCited by top-tier papers4
- Towards Understanding the Dynamics of Low-Rank AdaptationShu Ding, Yang Peng, Hangan Zhou, Xinyu Lu et al.ICML 2026 · 13 citations
- On the Convergence Rate of LoRA Gradient DescentSiqiao Mu, Diego KlabjanICML 2026 · 8 citations
- Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable DataHancheng Min, Zhihui Zhu, René VidalNeurIPS 2025 · 3 citations
- LoRA-One: One-Step Full Gradient Could Suffice for Fine-Tuning Large Language Models, Provably and EfficientlyYuanhe Zhang, Fanghui Liu, Yudong ChenICML 2025
Builds on11
- LoRA+: Efficient Low Rank Adaptation of Large ModelsSoufiane Hayou, Nikhil Ghosh, Bin YuICML 2024 · 388 citations
- The Expressive Power of Low-Rank AdaptationYuchen Zeng, Kangwook LeeICLR 2024 · 116 citations
- A Kernel-Based View of Language Model Fine-TuningSadhika Malladi, Alexander Wettig, Dingli Yu, Danqi Chen et al.ICML 2023 · 111 citations
- The Power of Scale for Parameter-Efficient Prompt TuningBrian Lester, Rami Al-Rfou, Noah ConstantEMNLP 2021 · 94 citations
- Weight decay induces low-rank attention layersSeijin Kobayashi, Yassir Akram, Johannes von OswaldNeurIPS 2024 · 41 citations
Related papers
- LoRA Training in the NTK Regime has No Spurious Local MinimaUijeong Jang, Jason D. Lee, Ernest K. RyuICML 2024 · 41 citations
- AdaRankGrad: Adaptive Gradient Rank and Moments for Memory-Efficient LLMs Training and Fine-TuningYehonathan Refael, Jonathan Svirsky, Boris Shustin, Wasim Huleihel et al.ICLR 2025
- Balanced LoRA: Removing Parameter Invariance to Accelerate ConvergenceValérie Castin, Kimia Nadjahi, Pierre Ablin, Gabriel PeyréICML 2026
- GeoLoRA: Geometric integration for parameter efficient fine-tuningSteffen Schotthöfer, Emanuele Zangrando, Gianluca Ceruti, Francesco Tudisco et al.ICLR 2025
- Stable-LoRA: Stabilizing Feature Learning of Low-Rank AdaptationYize Wu, Ke Gao, Ling Li, Yanjun WuICLR 2026 · 1 citation
