Computational Limits of Low-Rank Adaptation (LoRA) Fine-Tuning for Transformer Models
Jerry Yao-Chieh Hu, Maojiang Su, En-Jui Kuo, Zhao Song, Han Liu
Abstract
We study the computational limits of Low-Rank Adaptation (LoRA) for finetuning transformer-based models using fine-grained complexity theory. Our key observation is that the existence of low-rank decompositions within the gradient computation of LoRA adaptation leads to possible algorithmic speedup. This allows us to (i) identify a phase transition behavior of efficiency assuming the Strong Exponential Time Hypothesis (SETH), and (ii) prove the existence of almost linear algorithms by controlling the LoRA update computation term by term. For the former, we identify a sharp transition in the efficiency of all possible rank-r LoRA update algorithms for transformers, based on specific norms resulting from the multiplications of the input sequence X, pretrained weights W ⋆ , and adapter matrices αBA/r. Specifically, we derive a shared upper bound threshold for such norms, and show that efficient (sub-quadratic) approximation algorithms of LoRA exist only below this threshold. For the latter, we prove the existence of almost linear approximation algorithms for LoRA adaptation by utilizing the hierarchical low-rank structures of LoRA gradients and approximating the gradients with a series of chained low-rank approximations. To showcase our theory, we consider two practical scenarios: partial (e.g., only W V and W Q ) and full adaptations (e.g., W Q , W V , and W K ) of weights in attention heads. * Code is available on OpenReview; full version and future updates are on arXiv. Published as a conference paper at ICLR 2025 The hardness of LoRA's forward pass is trivially characterized by (Alman and Song, 2023). To see this, let X ∈ R L×d be input with length L, and W K , W Q , W V ∈ R d×d be attention weights, and with the inverse temperature β > 0 and Here, exp(•) is entry-wise exponential function, diag (•) converts a vector into a diagonal matrix with the entries of the vector, and 1 L is the length-L all ones vector. LoRA finetuning is given as Definition 1.1 (LoRA (Hu et al., 2021) ). Let W ∈ R b×a be any weight matrix in a pretrained model F , LoRA fine-tunes F through updating W with a low-rank decomposition W = W ⋆ + α r BA. Here, W ⋆ is the frozen pretrained weight. Only B ∈ R b×r and A ∈ R r×a are learnable (being update via gradient descent) with rank r < min(a, b) and tunable hyperparameter α ∈ R. Under the Strong Exponential Time Hypothesis (Hypothesis 1), Alman and Song (2023) state: Lemma 1.1 (Informal, (Alman and Song, 2023)). Fast (sub-quadratic) forward pass of transformer only exist when entries of K, Q, V are bounded by a constant B = Θ( √ log L). for some ϵ > 0, where ∥Z∥ ∞ := max i,j |Z ij |.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3e4bcdf6-3944-44c0-b76b-b7d75808bedfCited by top-tier papers5
- RHYTHM: Reasoning with Hierarchical Temporal Tokenization for Human MobilityHaoyu He, Haozheng Luo, Yan Chen, Qi (Cheems) WangNeurIPS 2025 · 7 citations
- Two Heads are Better than One: Simulating Large Transformers with Small OnesHantao Yu, Josh AlmanNeurIPS 2025 · 1 citation
- Fundamental Limits of Prompt Tuning Transformers: Universality, Capacity and EfficiencyJerry Yao-Chieh Hu, Wei-Po Wang, Ammar Gilani, Chenyang Li et al.ICLR 2025
- SumRA: Parameter Efficient Fine-tuning with Singular Value Decomposition and Summed Orthogonal BasisKwok Chin Yuen, Yongsen Zheng, Jia Qi Yip, Kwok-Yan Lam et al.ICLR 2026
- Fast and Low-Cost Genomic Foundation Models via Outlier RemovalHaozheng Luo, Chenghao Qiu, Maojiang Su, Zhihan Zhou et al.ICML 2025
Builds on28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide ResolutionEric Nguyen, Michael Poli, Marjan Faizi, Armin W. Thomas et al.NeurIPS 2023 · 574 citations
- LoRA+: Efficient Low Rank Adaptation of Large ModelsSoufiane Hayou, Nikhil Ghosh, Bin YuICML 2024 · 388 citations
Related papers
- The Expressive Power of Low-Rank AdaptationYuchen Zeng, Kangwook LeeICLR 2024 · 116 citations
- AdaRankGrad: Adaptive Gradient Rank and Moments for Memory-Efficient LLMs Training and Fine-TuningYehonathan Refael, Jonathan Svirsky, Boris Shustin, Wasim Huleihel et al.ICLR 2025
- On the Convergence Rate of LoRA Gradient DescentSiqiao Mu, Diego KlabjanICML 2026 · 8 citations
- GeoLoRA: Geometric integration for parameter efficient fine-tuningSteffen Schotthöfer, Emanuele Zangrando, Gianluca Ceruti, Francesco Tudisco et al.ICLR 2025
- GraLoRA: Granular Low-Rank Adaptation for Parameter-Efficient Fine-TuningYeonjoon Jung, Daehyun Ahn, Hyungjun Kim, Taesu Kim et al.NeurIPS 2025 · 11 citations
