GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
Jiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang, Anima Anandkumar, Yuandong Tian
Abstract
Training Large Language Models (LLMs) presents significant memory challenges, predominantly due to the growing size of weights and optimizer states. Common memory-reduction approaches, such as low-rank adaptation (LoRA), add a trainable low-rank matrix to the frozen pre-trained weight in each layer, reducing trainable parameters and optimizer states. However, such approaches typically underperform training with full-rank weights in both pre-training and fine-tuning stages since they limit the parameter search to a low-rank subspace and alter the training dynamics, and further, may require full-rank warm start. In this work, we propose Gradient Low-Rank Projection (GaLore), a training strategy that allows full-parameter learning but is more memory-efficient than common low-rank adaptation methods such as LoRA. Our approach reduces memory usage by up to 65.5% in optimizer states while maintaining both efficiency and performance for pre-training on LLaMA 1B and 7B architectures with C4 dataset with up to 19.7B tokens, and on fine-tuning RoBERTa on GLUE tasks. Our 8-bit GaLore further reduces optimizer memory by up to 82.5% and total training memory by 63.3%, compared to a BF16 baseline. Notably, we demonstrate, for the first time, the feasibility of pre-training a 7B model on consumer GPUs with 24GB memory (e.g., NVIDIA RTX 4090) without model parallel, checkpointing, or offloading strategies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b052d92e-adca-4c90-be63-179882b2f575Cited by top-tier papers186
- LoRA-GA: Low-Rank Adaptation with Gradient ApproximationShaowen Wang, Linxi Yu, Jian LiNeurIPS 2024 · 194 citations
- LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-TuningRui Pan, Xiang Liu, Shizhe Diao, Renjie Pi et al.NeurIPS 2024 · 124 citations
- Fantastic Pretraining Optimizers and Where to Find ThemKaiyue Wen, David Leo Wright Hall, Tengyu Ma, Percy LiangICLR 2026 · 92 citations
- The Curse of Depth in Large Language ModelsWenfang Sun, Xinyuan Song, Pengxiang Li, Lu Yin et al.NeurIPS 2025 · 62 citations
- Full Parameter Fine-tuning for Large Language Models with Limited ResourcesKai Lv, Yuqing Yang, Tengxiao Liu, Qipeng Guo et al.ACL 2024 · 61 citations
Builds on13
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
- 8-bit Optimizers via Block-wise QuantizationTim Dettmers, Mike Lewis, Sam Shleifer, Luke ZettlemoyerICLR 2022 · 457 citations
- ReLoRA: High-Rank Training Through Low-Rank UpdatesVladislav Lialin, Sherin Muckatira, Namrata Shivagunde, Anna RumshiskyICLR 2024 · 214 citations
Related papers
- Gradient Weight-normalized Low-rank Projection for Efficient LLM TrainingJia-Hong Huang, Yixian Shen, Hongyi Zhu, Stevan Rudinac et al.AAAI 2025 · 1 citation
- On the Optimization Landscape of Low Rank Adaptation Methods for Large Language ModelsXu-Hui Liu, Yali Du, Jun Wang, Yang YuICLR 2025
- Fira: Can We Achieve Full-rank Training of LLMs Under Low-rank Constraint?Xi Chen, Kaituo Feng, Changsheng Li, Xunhao Lai et al.NeurIPS 2025 · 48 citations
- COAP: Memory-Efficient Training with Correlation-Aware Gradient ProjectionJinqi Xiao, Shen Sang, Tiancheng Zhi, Jing Liu et al.CVPR 2025
- FRUGAL: Memory-Efficient Optimization by Reducing State Overhead for Scalable TrainingPhilip Zmushko, Aleksandr Beznosikov, Martin Takác, Samuel HorváthICML 2025
