STEP: Learning N: M Structured Sparsity Masks from Scratch with Precondition
Yucheng Lu, Shivani Agrawal, Suvinay Subramanian, Oleg Rybakov, Christopher De Sa, Amir Yazdanbakhsh
Abstract
Recent innovations on hardware (e.g. Nvidia A100) have motivated learning N:M structured sparsity masks from scratch for fast model inference. However, state-of-the-art learning recipes in this regime (e.g. SR-STE) are proposed for non-adaptive optimizers like momentum SGD, while incurring non-trivial accuracy drop for Adam-trained models like attention-based LLMs. In this paper, we first demonstrate such gap origins from poorly estimated second moment (i.e. variance) in Adam states given by the masked weights. We conjecture that learning N:M masks with Adam should take the critical regime of variance estimation into account. In light of this, we propose STEP, an Adam-aware recipe that learns N:M masks with two phases: first, STEP calculates a reliable variance estimate (precondition phase) and subsequently, the variance remains fixed and is used as a precondition to learn N:M masks (mask-learning phase). STEP automatically identifies the switching point of two phases by dynamically sampling variance changes over the training trajectory and testing the sample concentration. Empirically, we evaluate STEP and other baselines such as ASP and SR-STE on multiple tasks including CIFAR classification, machine translation and LLM fine-tuning (BERT-Base, GPT-2). We show STEP mitigates the accuracy drop of baseline recipes and is robust to aggressive structured sparsity ratios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0a7d1f65-6925-41b5-9f3a-6ee32bbafba3Cited by top-tier papers15
- MaskLLM: Learnable Semi-Structured Sparsity for Large Language ModelsGongfan Fang, Hongxu Yin, Saurav Muralidharan, Greg Heinrich et al.NeurIPS 2024 · 72 citations
- Scaling Laws for Sparsely-Connected Foundation ModelsElias Frantar, Carlos Riquelme Ruiz, Neil Houlsby, Dan Alistarh et al.ICLR 2024 · 48 citations
- S-STE: Continuous Pruning Function for Efficient 2: 4 Sparse Pre-trainingYuezhou Hu, Jun Zhu, Jianfei ChenNeurIPS 2024 · 14 citations
- TB-STC: Transposable Block-wise N: M Structured Sparse Tensor CoreJun Liu, Shulin Zeng, Junbo Zhao, Li Ding et al.HPCA 2025 · 9 citations
- Cut Less, Fold More: Model Compression through the Lens of Projection GeometryOlga Saukh, Dong Wang, Haris Sikic, Yun Cheng et al.ICLR 2026 · 4 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
- Why are Adaptive Methods Good for Attention Models?Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim et al.NeurIPS 2020 · 397 citations
Related papers
- Learning N: M Fine-grained Structured Sparse Neural Networks From ScratchAojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu et al.ICLR 2021 · 301 citations
- Adaptive Preconditioners Trigger Loss Spikes in AdamZhiwei Bai, Zhangchen Zhou, Jiajie Zhao, Xiaolong Li et al.ICML 2026 · 9 citations
- MoMo: Momentum Models for Adaptive Learning RatesFabian Schaipp, Ruben Ohana, Michael Eickenberg, Aaron Defazio et al.ICML 2024 · 21 citations
- Understanding and improving Shampoo and SOAP via Kullback-Leibler MinimizationWu Lin, Scott C. Lowe, Felix Dangel, Runa Eschenhagen et al.ICLR 2026 · 15 citations
- Minimum Variance Unbiased N: M Sparsity for the Neural GradientsBrian Chmiel, Itay Hubara, Ron Banner, Daniel SoudryICLR 2023
