Efficient Unstructured Pruning of Mamba State-Space Models for Resource-Constrained Environments
Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma
Abstract
As AI deployment shifts to edge devices, efficient sequence modeling becomes critical.State-space models (SSMs), particularly Mamba, rival Transformers with linear-time complexity and strong performance across tasks, yet their large parameter counts hinder resource-constrained use.We propose a novel unstructured pruning framework tailored for Mamba, achieving up to 70% parameter reduction with only 3-9% performance loss.Unlike Transformer-focused pruning, our approach leverages Mamba's recurrent dynamics through: (1) pruning based on weight and gradient importance to preserve critical parameters, (2) a gradual pruning process to ensure model stability, and (3) a global strategy optimizing parameter allocation across the model.Extensive experiments on WikiText-103, Long Range Arena, and ETT benchmarks show significant efficiency gains, with 1.77 faster inference and 46% less memory.Our component analysis reveals Mamba's robustness, enabling practical deployment while requiring careful use to avoid biases in sensitive applications.ronments.State-space models (SSMs) (Gu et al., 2020a(Gu et al., , 2021;; Gupta et al., 2022) offer a promising alternative with linear-time complexity while effectively modeling long-range dependencies.The Mamba architecture (Gu and Dao, 2023) distinguishes itself through its selective mechanism that dynamically controls information flow based * Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9df7dff-2e49-4534-88b4-9826e282aa4dCited by top-tier papers3
- HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL ConferencesYusuke Sakai, Hidetaka Kamigaito, Taro WatanabeACL 2026 · 18 citations
- UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMsHung-Yueh Chiang, Chi-Chih Chang, Yu-Chen Lu, Chien-Yu Lin et al.ICLR 2026 · 6 citations
- Beyond Variance: Knowledge-Aware LLM Compression via Fisher-Aligned Subspace DiagnosticsIbne Farabi Shihab, Sanjeda Akter, Anuj SharmaACL 2026 · 4 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
Related papers
- SparseSSM: Efficient Selective Structured State Space Models Can Be Pruned in One-ShotKaiwen TUO, Huan WangICML 2026 · 6 citations
- Rethinking Token Reduction for State Space ModelsZheng Zhan, Yushu Wu, Zhenglun Kong, Changdi Yang et al.EMNLP 2024 · 4 citations
- TransMamba: A Sequence-Level Hybrid Transformer-Mamba Language ModelYixing Li, Ruobing Xie, Zhen Yang, Xingwu Sun et al.AAAI 2026 · 3 citations
- Quamba: A Post-Training Quantization Recipe for Selective State Space ModelsHung-Yueh Chiang, Chi-Chih Chang, Natalia Frumkin, Kai-Chiang Wu et al.ICLR 2025
- MambaPEFT: Exploring Parameter-Efficient Fine-Tuning for MambaMasakazu Yoshimura, Teruaki Hayashi, Yota MaedaICLR 2025
