GRASS: Compute Efficient Low-Memory LLM Training with Structured Sparse Gradients
Aashiq Muhamed, Oscar Li, David P. Woodruff, Mona T. Diab, Virginia Smith
Abstract
Large language model (LLM) training and finetuning are often bottlenecked by limited GPU memory. While existing projection-based optimization methods address this by projecting gradients into a lower-dimensional subspace to reduce optimizer state memory, they typically rely on dense projection matrices, which can introduce computational and memory overheads. In this work, we propose GRASS (GRAdient Stuctured Sparsification), a novel approach that leverages sparse projections to transform gradients into structured sparse updates. This design not only significantly reduces memory usage for optimizer states but also minimizes gradient memory footprint, computation, and communication costs, leading to substantial throughput improvements. Extensive experiments on pretraining and finetuning tasks demonstrate that GRASS achieves competitive performance to full-rank training and existing projection-based methods. Notably, GRASS enables half-precision pretraining of a 13B parameter LLaMA model on a single 40GB A100 GPU-a feat infeasible for previous methodsand yields up to a 2× throughput improvement on an 8-GPU system. Code is released here 1 . P ← compute P (∇L(W (t) )) ▷ P ∈ R m×r 8: // [Optional] Update optimizer state 9: S (t) ← update_state(S (t) ) 10: end if 11: S (t+1) , ∆ (t+1) ← opt.update(S (t) , GC ) 13: ▷ Apply update 14: end for Algorithm 2 MeSO Implementations FLORA Compute dense P : Sample Pij i.i.d. from N (0, 1/r). Update_state: Updates momentum as P (t+1) P ⊤ (t) S (t) . Compute GC : Computes GC using dense matmul. Apply update: Updates full W after dense matmul. GALORE Compute dense P : Top-r left singular vectors of grad GW . Update_state: Maintains optimizer state. Compute GC : Computes GC using dense matmul. Apply update: Updates full W after a dense matmul. GRASS (ours) Compute sparse P : Computes the selection matrix B and the diagonal scaling matrix ρ based on row norms of GW . Update_state: Resets S (t) to zero as necessary. Compute GC : Uses matrix associativity and sparse matmul. Apply update: Sparse update W after sparse matmul. structured sparse matrices for P , demonstrating their advantages in memory, computation, and communication efficiency across both pretraining and finetuning. Our main contributions include:
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 102d1e35-0727-48c9-aabe-e0ef93e30a17Cited by top-tier papers9
- SubTrack++ : Gradient Subspace Tracking for Scalable LLM TrainingSahar Rajabi, Nayeema Nonta, Sirisha RambhatlaNeurIPS 2025 · 19 citations
- Memory-Efficient LLM Pretraining via Minimalist Optimizer DesignAthanasios Glentis, Jiaxiang Li, Andi Han, Mingyi HongICML 2026 · 9 citations
- Breaking the Frozen Subspace: Importance Sampling for Low-Rank Optimization in LLM PretrainingHaochen Zhang, Junze Yin, Guanchu Wang, Zirui Liu et al.NeurIPS 2025 · 7 citations
- QKV Projections Require a Fraction of Their MemoryMalik Khalaf, Yara Shamshoum, Nitzan Hodos, Yuval Sieradzki et al.ICLR 2026 · 3 citations
- INSTANT: Compressing Gradients and Activations for Resource-Efficient TrainingTuan-Kiet Doan, Trung-Hieu Tran, Enzo Tartaglione, Nikola Simidjievski et al.ICLR 2026
Builds on9
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- 8-bit Optimizers via Block-wise QuantizationTim Dettmers, Mike Lewis, Sam Shleifer, Luke ZettlemoyerICLR 2022 · 457 citations
Related papers
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang et al.ICML 2024 · 433 citations
- FRUGAL: Memory-Efficient Optimization by Reducing State Overhead for Scalable TrainingPhilip Zmushko, Aleksandr Beznosikov, Martin Takác, Samuel HorváthICML 2025
- A Memory Efficient Randomized Subspace Optimization Method for Training Large Language ModelsYiming Chen, Yuan Zhang, Yin Liu, Kun Yuan et al.ICML 2025
- Full Parameter Fine-tuning for Large Language Models with Limited ResourcesKai Lv, Yuqing Yang, Tengxiao Liu, Qipeng Guo et al.ACL 2024 · 61 citations
- SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM TrainingTianjin Huang, Ziquan Zhu, Gaojie Jin, Lu Liu et al.ICLR 2025
