Sparse Learning for State Space Models on Mobile
Xuan Shen, Hangyu Zheng, Yifan Gong, Zhenglun Kong, Changdi Yang, Zheng Zhan, Yushu Wu, Xue Lin, Yanzhi Wang, Pu Zhao, Wei Niu
Abstract
Transformer models have been widely investigated in different domains by providing long-range dependency handling and global contextual awareness, driving the development of popular AI applications such as ChatGPT, Gemini, and Alexa. State Space Models (SSMs) have emerged as strong contenders in the field of sequential modeling, challenging the dominance of Transformers. SSMs incorporate a selective mechanism that allows for dynamic parameter adjustment based on input data, enhancing their performance. However, this mechanism also comes with increasing computational complexity and bandwidth demands, posing challenges for deployment on resource-constraint mobile devices. To address these challenges without sacrificing the accuracy of the selective mechanism, we propose a sparse learning framework that integrates architecture-aware compiler optimizations. We introduce an end-to-end solution-C n 4 kernel sparsity, which prunes n elements from every four contiguous weights, and develop a compiler-based acceleration solution to ensure execution efficiency for this sparsity on mobile devices. Based on the kernel sparsity, our framework generates optimized sparse models targeting specific sparsity or latency requirements for various model sizes. We further leverage pruned weights to compensate for the remaining weights, enhancing downstream task performance. For practical hardware acceleration, we propose C n 4 -specific optimizations combined with a layout transformation elimination strategy. This approach mitigates inefficiencies arising from fine-grained pruning in linear layers and improves performance across other operations. Experimental results demonstrate that our method achieves superior task performance compared to other semi-structured pruning methods and achieves up-to 7→ speedup compared to llama.cpp framework on mobile devices. 1. We design a special kernel C n 4 and with a set of comprehensive compiler optimizations, including C n 4 -specific optimizations and layout transformation elimination strategy on mobile devices. 2. We propose the sparsity-oriented and/or latency-oriented sparse learning framework to explore the optimal pruning strategy with the proposed kernels for Mamba models. 3. We propose the weight compensation algorithm for the rectification of the sparse model weights by calibrating with only 128 samples, thereby further enhancing the model effectiveness. 4. Experiments show that our framework can achieve better task performance than other semistructure pruning methods and achieve pratical on-device speedup up to 7→ compared to llama.cpp.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03c4e7f2-04d6-4209-8443-6bcd08888d05Cited by top-tier papers2
- Efficient Reasoning with Hidden ThinkingXuan Shen, Yizhou Wang, Yufa Zhou, Xiangxi Shi et al.ICML 2026 · 56 citations
- QuartDepth: Post-Training Quantization for Real-Time Depth Estimation on the EdgeXuan Shen, Weize Ma, Jing Liu, Changdi Yang et al.CVPR 2025
Builds on25
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
Related papers
- UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMsHung-Yueh Chiang, Chi-Chih Chang, Yu-Chen Lu, Chien-Yu Lin et al.ICLR 2026 · 6 citations
- Efficient Unstructured Pruning of Mamba State-Space Models for Resource-Constrained EnvironmentsIbne Farabi Shihab, Sanjeda Akter, Anuj SharmaEMNLP 2025 · 1 citation
- Can Mamba Learn How To Learn? A Comparative Study on In-Context Learning TasksJongho Park, Jaeseung Park, Zheyang Xiong, Nayoung Lee et al.ICML 2024 · 124 citations
- NxMTransformer: Semi-Structured Sparsification for Natural Language Understanding via ADMMConnor Holmes, Minjia Zhang, Yuxiong He, Bo WuNeurIPS 2021 · 29 citations
- Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionJingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo et al.ACL 2025 · 334 citations
