AdaRankGrad: Adaptive Gradient Rank and Moments for Memory-Efficient LLMs Training and Fine-Tuning
Yehonathan Refael, Jonathan Svirsky, Boris Shustin, Wasim Huleihel, Ofir Lindenbaum
Abstract
Training and fine-tuning large language models (LLMs) come with challenges related to memory and computational requirements due to the increasing size of the model weights and the optimizer states. Various techniques have been developed to tackle these challenges, such as low-rank adaptation (LoRA), which involves introducing a parallel trainable low-rank matrix to the fixed pre-trained weights at each layer. However, these methods often fall short compared to the full-rank weight training approach, as they restrict the parameter search to a low-rank subspace. This limitation can disrupt training dynamics and require a full-rank warm start to mitigate the impact. In this paper, we introduce a new method inspired by a phenomenon we formally prove: as training progresses, the rank of the estimated layer gradients gradually decreases, and asymptotically approaches rank one. Leveraging this, our approach involves adaptively reducing the rank of the gradients during Adam optimization steps, using an efficient online-updating low-rank projections rule. We further present a randomized SVD scheme for efficiently finding the projection matrix. Our technique enables full-parameter fine-tuning with adaptive low-rank gradient updates, significantly reducing overall memory requirements during training compared to state-of-the-art methods while improving model performance in both pretraining and fine-tuning. Finally, we provide a convergence analysis of our method and demonstrate its merits for training and fine-tuning language and biological foundation models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3009f36e-016b-4809-a69f-9cee3e97b7daCited by top-tier papers4
- SUMO: Subspace-Aware Moment-Orthogonalization for Accelerating Memory-Efficient LLM TrainingYehonathan Refael, Guy Smorodinsky, Tom Tirer, Ofir LindenbaumNeurIPS 2025 · 17 citations
- ICR-RL: Deep Reinforcement Learning via In-Context-RegressionDavid Schiff, Ofir Lindenbaum, Yonathan EfroniICML 2026 · 5 citations
- Beyond A Single AI Cluster: A Survey of Decentralized LLM TrainingHaotian Dong, Jingyan Jiang, Rongwei Lu, Jiajun Luo et al.EMNLP 2025 · 2 citations
- No Prior, No Leakage: Revisiting Reconstruction Attacks in Trained Neural NetworksYehonathan Refael, Guy Smorodinsky, Ofir Lindenbaum, Itay SafranICLR 2026 · 1 citation
Builds on8
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang et al.ICML 2024 · 433 citations
- ReLoRA: High-Rank Training Through Low-Rank UpdatesVladislav Lialin, Sherin Muckatira, Namrata Shivagunde, Anna RumshiskyICLR 2024 · 214 citations
- Memory-Efficient Fine-Tuning of Compressed Large Language Models via sub-4-bit Integer QuantizationJeonghoon Kim, Jung Hyun Lee, Sungdong Kim, Joonsuk Park et al.NeurIPS 2023 · 157 citations
Related papers
- ScaLoRA: Optimally Scaled Low-Rank Adaptation for Efficient High-Rank Fine-TuningYilang Zhang, Xiaodong Yang, Yiwei Cai, Georgios B. GiannakisICML 2026 · 1 citation
- LoFT: Low-Rank Adaptation That Behaves Like Full Fine-TuningNurbek Tastan, Stefanos Laskaridis, Martin Takác, Karthik Nandakumar et al.ICLR 2026 · 16 citations
- AltLoRA: Towards Better Gradient Approximation in Low-Rank Adaptation with Alternating ProjectionsXin Yu, Yujia Wang, Jinghui Chen, Lingzhou XueNeurIPS 2025 · 8 citations
- LoRA-Pro: Are Low-Rank Adapters Properly Optimized?Zhengbo Wang, Jian Liang, Ran He, Zilei Wang et al.ICLR 2025
- RandLoRA: Full rank parameter-efficient fine-tuning of large modelsPaul Albert, Frederic Z. Zhang, Hemanth Saratchandran, Cristian Rodriguez Opazo et al.ICLR 2025
