GRASS: Compute Efficient Low-Memory LLM Training with Structured Sparse Gradients
Aashiq Muhamed, Oscar Li, David P. Woodruff, Mona T. Diab, Virginia Smith
摘要
Large language model (LLM) training and finetuning are often bottlenecked by limited GPU memory. While existing projection-based optimization methods address this by projecting gradients into a lower-dimensional subspace to reduce optimizer state memory, they typically rely on dense projection matrices, which can introduce computational and memory overheads. In this work, we propose GRASS (GRAdient Stuctured Sparsification), a novel approach that leverages sparse projections to transform gradients into structured sparse updates. This design not only significantly reduces memory usage for optimizer states but also minimizes gradient memory footprint, computation, and communication costs, leading to substantial throughput improvements. Extensive experiments on pretraining and finetuning tasks demonstrate that GRASS achieves competitive performance to full-rank training and existing projection-based methods. Notably, GRASS enables half-precision pretraining of a 13B parameter LLaMA model on a single 40GB A100 GPU-a feat infeasible for previous methodsand yields up to a 2× throughput improvement on an 8-GPU system. Code is released here 1 . P ← compute P (∇L(W (t) )) ▷ P ∈ R m×r 8: // [Optional] Update optimizer state 9: S (t) ← update_state(S (t) ) 10: end if 11: S (t+1) , ∆ (t+1) ← opt.update(S (t) , GC ) 13: ▷ Apply update 14: end for Algorithm 2 MeSO Implementations FLORA Compute dense P : Sample Pij i.i.d. from N (0, 1/r). Update_state: Updates momentum as P (t+1) P ⊤ (t) S (t) . Compute GC : Computes GC using dense matmul. Apply update: Updates full W after dense matmul. GALORE Compute dense P : Top-r left singular vectors of grad GW . Update_state: Maintains optimizer state. Compute GC : Computes GC using dense matmul. Apply update: Updates full W after a dense matmul. GRASS (ours) Compute sparse P : Computes the selection matrix B and the diagonal scaling matrix ρ based on row norms of GW . Update_state: Resets S (t) to zero as necessary. Compute GC : Uses matrix associativity and sparse matmul. Apply update: Sparse update W after sparse matmul. structured sparse matrices for P , demonstrating their advantages in memory, computation, and communication efficiency across both pretraining and finetuning. Our main contributions include:
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- SubTrack++ : Gradient Subspace Tracking for Scalable LLM TrainingSahar Rajabi, Nayeema Nonta, Sirisha RambhatlaNeurIPS 2025 · 被引用 19 次
- Memory-Efficient LLM Pretraining via Minimalist Optimizer DesignAthanasios Glentis, Jiaxiang Li, Andi Han, Mingyi HongICML 2026 · 被引用 9 次
- Breaking the Frozen Subspace: Importance Sampling for Low-Rank Optimization in LLM PretrainingHaochen Zhang, Junze Yin, Guanchu Wang, Zirui Liu 等NeurIPS 2025 · 被引用 7 次
- QKV Projections Require a Fraction of Their MemoryMalik Khalaf, Yara Shamshoum, Nitzan Hodos, Yuval Sieradzki 等ICLR 2026 · 被引用 3 次
- INSTANT: Compressing Gradients and Activations for Resource-Efficient TrainingTuan-Kiet Doan, Trung-Hieu Tran, Enzo Tartaglione, Nikola Simidjievski 等ICLR 2026
它引用的顶会 Paper9
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- 8-bit Optimizers via Block-wise QuantizationTim Dettmers, Mike Lewis, Sam Shleifer, Luke ZettlemoyerICLR 2022 · 被引用 457 次
相关 Paper
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang 等ICML 2024 · 被引用 433 次
- FRUGAL: Memory-Efficient Optimization by Reducing State Overhead for Scalable TrainingPhilip Zmushko, Aleksandr Beznosikov, Martin Takác, Samuel HorváthICML 2025
- A Memory Efficient Randomized Subspace Optimization Method for Training Large Language ModelsYiming Chen, Yuan Zhang, Yin Liu, Kun Yuan 等ICML 2025
- Full Parameter Fine-tuning for Large Language Models with Limited ResourcesKai Lv, Yuqing Yang, Tengxiao Liu, Qipeng Guo 等ACL 2024 · 被引用 61 次
- SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM TrainingTianjin Huang, Ziquan Zhu, Gaojie Jin, Lu Liu 等ICLR 2025
