Memory-Efficient LLM Training with Dynamic Sparsity: From Stability to Practical Scaling
Qiao Xiao, Boqian Wu, Patrik Okanovic, Tomasz Sternal, Maurice Keulen, Elena Mocanu, Mykola Pechenizkiy, Decebal Constantin Mocanu, Torsten Hoefler
Abstract
Dynamic Sparse Training (DST) offers a promising paradigm for improving the training and inference efficiency of deep neural networks; however, we find that in large language model training, DST can suffer from optimization instability, manifested as loss spikes after topology updates. In this work, we show that the naive use of standard Adam-based optimizers leads to a cold-start issue for newly regrown parameters, resulting in excessively large updates and disrupted training dynamics. To address this issue, we propose Sparse Memory-Efficient Training (SMET), which stabilizes DST with optimizer warm-up and improves training progress through density-aware learning-rate scaling. SMET further reduces memory consumption by storing gradients and optimizer states only for active parameters. We provide a theoretical analysis of the update behaviors under SMET, showing improved optimization stability. Extensive experiments demonstrate that SMET enables stable, scalable, and memory-efficient sparse pre-training of LLMs, paving the way for sparse training as a practical alternative to dense training. Our code is publicly available at: https://github.com/QiaoXiao7282/SMET.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0de2ff50-9d61-461c-a2e5-ee042edf8cd2Builds on20
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 884 citations
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang et al.ICML 2024 · 433 citations
- Chasing Sparsity in Vision Transformers: An End-to-End ExplorationTianlong Chen, Yu Cheng, Zhe Gan, Lu Yuan et al.NeurIPS 2021 · 295 citations
Related papers
- SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM TrainingTianjin Huang, Ziquan Zhu, Gaojie Jin, Lu Liu et al.ICLR 2025
- LDAdam: Adaptive Optimization from Low-Dimensional Gradient StatisticsThomas Robert, Mher Safaryan, Ionut-Vlad Modoranu, Dan AlistarhICLR 2025
- Dynamic Sparse No Training: Training-Free Fine-tuning for Sparse LLMsYuxin Zhang, Lirui Zhao, Mingbao Lin, Yunyun Sun et al.ICLR 2024 · 78 citations
- Sparse MeZO: Less Parameters for Better Performance in Zeroth-Order LLM Fine-TuningYong Liu, Zirui Zhu, Chaoyu Gong, Minhao Cheng et al.NeurIPS 2025 · 66 citations
- Pruning Large Language Models with Semi-Structural Adaptive Sparse TrainingWeiyu Huang, Yuezhou Hu, Guohao Jian, Jun Zhu et al.AAAI 2025 · 25 citations
