Computational Limits of Low-Rank Adaptation (LoRA) Fine-Tuning for Transformer Models
Jerry Yao-Chieh Hu, Maojiang Su, En-Jui Kuo, Zhao Song, Han Liu
摘要
We study the computational limits of Low-Rank Adaptation (LoRA) for finetuning transformer-based models using fine-grained complexity theory. Our key observation is that the existence of low-rank decompositions within the gradient computation of LoRA adaptation leads to possible algorithmic speedup. This allows us to (i) identify a phase transition behavior of efficiency assuming the Strong Exponential Time Hypothesis (SETH), and (ii) prove the existence of almost linear algorithms by controlling the LoRA update computation term by term. For the former, we identify a sharp transition in the efficiency of all possible rank-r LoRA update algorithms for transformers, based on specific norms resulting from the multiplications of the input sequence X, pretrained weights W ⋆ , and adapter matrices αBA/r. Specifically, we derive a shared upper bound threshold for such norms, and show that efficient (sub-quadratic) approximation algorithms of LoRA exist only below this threshold. For the latter, we prove the existence of almost linear approximation algorithms for LoRA adaptation by utilizing the hierarchical low-rank structures of LoRA gradients and approximating the gradients with a series of chained low-rank approximations. To showcase our theory, we consider two practical scenarios: partial (e.g., only W V and W Q ) and full adaptations (e.g., W Q , W V , and W K ) of weights in attention heads. * Code is available on OpenReview; full version and future updates are on arXiv. Published as a conference paper at ICLR 2025 The hardness of LoRA's forward pass is trivially characterized by (Alman and Song, 2023). To see this, let X ∈ R L×d be input with length L, and W K , W Q , W V ∈ R d×d be attention weights, and with the inverse temperature β > 0 and Here, exp(•) is entry-wise exponential function, diag (•) converts a vector into a diagonal matrix with the entries of the vector, and 1 L is the length-L all ones vector. LoRA finetuning is given as Definition 1.1 (LoRA (Hu et al., 2021) ). Let W ∈ R b×a be any weight matrix in a pretrained model F , LoRA fine-tunes F through updating W with a low-rank decomposition W = W ⋆ + α r BA. Here, W ⋆ is the frozen pretrained weight. Only B ∈ R b×r and A ∈ R r×a are learnable (being update via gradient descent) with rank r < min(a, b) and tunable hyperparameter α ∈ R. Under the Strong Exponential Time Hypothesis (Hypothesis 1), Alman and Song (2023) state: Lemma 1.1 (Informal, (Alman and Song, 2023)). Fast (sub-quadratic) forward pass of transformer only exist when entries of K, Q, V are bounded by a constant B = Θ( √ log L). for some ϵ > 0, where ∥Z∥ ∞ := max i,j |Z ij |.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- RHYTHM: Reasoning with Hierarchical Temporal Tokenization for Human MobilityHaoyu He, Haozheng Luo, Yan Chen, Qi (Cheems) WangNeurIPS 2025 · 被引用 7 次
- Two Heads are Better than One: Simulating Large Transformers with Small OnesHantao Yu, Josh AlmanNeurIPS 2025 · 被引用 1 次
- Fundamental Limits of Prompt Tuning Transformers: Universality, Capacity and EfficiencyJerry Yao-Chieh Hu, Wei-Po Wang, Ammar Gilani, Chenyang Li 等ICLR 2025
- SumRA: Parameter Efficient Fine-tuning with Singular Value Decomposition and Summed Orthogonal BasisKwok Chin Yuen, Yongsen Zheng, Jia Qi Yip, Kwok-Yan Lam 等ICLR 2026
- Fast and Low-Cost Genomic Foundation Models via Outlier RemovalHaozheng Luo, Chenghao Qiu, Maojiang Su, Zhihan Zhou 等ICML 2025
它引用的顶会 Paper28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng 等ICML 2020 · 被引用 1,388 次
- HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide ResolutionEric Nguyen, Michael Poli, Marjan Faizi, Armin W. Thomas 等NeurIPS 2023 · 被引用 574 次
- LoRA+: Efficient Low Rank Adaptation of Large ModelsSoufiane Hayou, Nikhil Ghosh, Bin YuICML 2024 · 被引用 388 次
相关 Paper
- The Expressive Power of Low-Rank AdaptationYuchen Zeng, Kangwook LeeICLR 2024 · 被引用 116 次
- AdaRankGrad: Adaptive Gradient Rank and Moments for Memory-Efficient LLMs Training and Fine-TuningYehonathan Refael, Jonathan Svirsky, Boris Shustin, Wasim Huleihel 等ICLR 2025
- On the Convergence Rate of LoRA Gradient DescentSiqiao Mu, Diego KlabjanICML 2026 · 被引用 8 次
- GeoLoRA: Geometric integration for parameter efficient fine-tuningSteffen Schotthöfer, Emanuele Zangrando, Gianluca Ceruti, Francesco Tudisco 等ICLR 2025
- GraLoRA: Granular Low-Rank Adaptation for Parameter-Efficient Fine-TuningYeonjoon Jung, Daehyun Ahn, Hyungjun Kim, Taesu Kim 等NeurIPS 2025 · 被引用 11 次
