Efficient Unstructured Pruning of Mamba State-Space Models for Resource-Constrained Environments
Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma
摘要
As AI deployment shifts to edge devices, efficient sequence modeling becomes critical.State-space models (SSMs), particularly Mamba, rival Transformers with linear-time complexity and strong performance across tasks, yet their large parameter counts hinder resource-constrained use.We propose a novel unstructured pruning framework tailored for Mamba, achieving up to 70% parameter reduction with only 3-9% performance loss.Unlike Transformer-focused pruning, our approach leverages Mamba's recurrent dynamics through: (1) pruning based on weight and gradient importance to preserve critical parameters, (2) a gradual pruning process to ensure model stability, and (3) a global strategy optimizing parameter allocation across the model.Extensive experiments on WikiText-103, Long Range Arena, and ETT benchmarks show significant efficiency gains, with 1.77 faster inference and 46% less memory.Our component analysis reveals Mamba's robustness, enabling practical deployment while requiring careful use to avoid biases in sensitive applications.ronments.State-space models (SSMs) (Gu et al., 2020a(Gu et al., , 2021;; Gupta et al., 2022) offer a promising alternative with linear-time complexity while effectively modeling long-range dependencies.The Mamba architecture (Gu and Dao, 2023) distinguishes itself through its selective mechanism that dynamically controls information flow based * Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL ConferencesYusuke Sakai, Hidetaka Kamigaito, Taro WatanabeACL 2026 · 被引用 18 次
- UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMsHung-Yueh Chiang, Chi-Chih Chang, Yu-Chen Lu, Chien-Yu Lin 等ICLR 2026 · 被引用 6 次
- Beyond Variance: Knowledge-Aware LLM Compression via Fisher-Aligned Subspace DiagnosticsIbne Farabi Shihab, Sanjeda Akter, Anuj SharmaACL 2026 · 被引用 4 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang 等AAAI 2021 · 被引用 7,289 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
相关 Paper
- SparseSSM: Efficient Selective Structured State Space Models Can Be Pruned in One-ShotKaiwen TUO, Huan WangICML 2026 · 被引用 6 次
- Rethinking Token Reduction for State Space ModelsZheng Zhan, Yushu Wu, Zhenglun Kong, Changdi Yang 等EMNLP 2024 · 被引用 4 次
- TransMamba: A Sequence-Level Hybrid Transformer-Mamba Language ModelYixing Li, Ruobing Xie, Zhen Yang, Xingwu Sun 等AAAI 2026 · 被引用 3 次
- Quamba: A Post-Training Quantization Recipe for Selective State Space ModelsHung-Yueh Chiang, Chi-Chih Chang, Natalia Frumkin, Kai-Chiang Wu 等ICLR 2025
- MambaPEFT: Exploring Parameter-Efficient Fine-Tuning for MambaMasakazu Yoshimura, Teruaki Hayashi, Yota MaedaICLR 2025
