Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models
Hung-Yueh Chiang, Chi-Chih Chang, Natalia Frumkin, Kai-Chiang Wu, Mohamed S. Abdelfattah, Diana Marculescu
摘要
Despite State Space Models (SSMs) are emerging as an efficient alternative to Transformers, depolying SSMs on both cloud and edge devices is challenging due to the limited resources. Model quantization reduces model size and leverages hardware acceleration, and recent efforts on SSM quantization have focused on optimizing a particular model or bit-width. However, distinct bitwidths are essential for different scenarios, like W4A8 for boosting large-batch decoding speed, and W4A16 for enhancing generation speed in short-prompt single-user applications. We present Quamba2, compatible with W8A8, W4A8, and W4A16 for both Mamba1 and Mamba2 backbones, addressing the growing demand for SSM deployment. Based on channel order preserving and activation persistence of SSMs, we propose an offline approach to quantize inputs of the linear recurrence in 8-bit by sorting and clustering for input x, combined with a per-state-group quantization for input-dependent parameters B and C. To ensure compute-invariance in the SSM output, we rearrange weights offline according to the clustering sequence. We show that Quamba2-8B outperforms two state-of-the-art SSM quantization methods and delivers 1.3× and 3× speed-ups in the pre-filling and generation stages, respectively, while offering 4× memory reduction with only a 1.6% average accuracy drop. The evaluation on MMLU shows the generalizability and robustness of our framework. The code and quantized models are released at the link.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMsHung-Yueh Chiang, Chi-Chih Chang, Yu-Chen Lu, Chien-Yu Lin 等ICLR 2026 · 被引用 6 次
- SSDi8: Accurate and Efficient 8-bit Quantization for State Space DualityHyunwoo Kim, Byoungchan Ko, Minseok Kang, Minwoo Kim 等ICLR 2026 · 被引用 3 次
- Efficient Hybrid Language Model Compression through Group-Aware SSM PruningAli Taghibakhshi, Sharath Turuvekere Sreenivas, Saurav Muralidharan, Marcin Chochowski 等NeurIPS 2025 · 被引用 1 次
它引用的顶会 Paper13
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang 等ICML 2024 · 被引用 1,725 次
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 被引用 1,407 次
相关 Paper
- Quamba: A Post-Training Quantization Recipe for Selective State Space ModelsHung-Yueh Chiang, Chi-Chih Chang, Natalia Frumkin, Kai-Chiang Wu 等ICLR 2025
- ViM-VQ: Efficient Post-Training Vector Quantization for Visual MambaJuncan Deng, Shuaiting Li, Zeyu Wang, Kedong Xu 等ICCV 2025 · 被引用 2 次
- MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation MethodsZukang Xu, Yuxuan Yue, Xing Hu, Dawei Yang 等ICLR 2025
- Efficient Unstructured Pruning of Mamba State-Space Models for Resource-Constrained EnvironmentsIbne Farabi Shihab, Sanjeda Akter, Anuj SharmaEMNLP 2025 · 被引用 1 次
- Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic ModelsAviv Bick, Kevin Y. Li, Eric P. Xing, J. Zico Kolter 等NeurIPS 2024 · 被引用 78 次
