Efficient Hybrid Language Model Compression through Group-Aware SSM Pruning
Ali Taghibakhshi, Sharath Turuvekere Sreenivas, Saurav Muralidharan, Marcin Chochowski, Yashaswi Karnati, Raviraj Joshi, Ameya Mahabaleshwarkar, Zijia Chen, Yoshi Suhara, Oluwatobi Olabiyi, Daniel Korzekwa, Mostofa Patwary
摘要
Hybrid language models that combine Attention and State Space Models (SSMs) have been shown to achieve state-of-the-art accuracy and runtime performance. Recent work has also demonstrated that applying pruning and distillation to Attentiononly models yields smaller, more accurate models at a fraction of the training cost. In this work, we explore the effectiveness of compressing Hybrid architectures. To this end, we introduce a novel group-aware pruning method for Mamba layers that preserves the structural integrity of SSM blocks and their sequence modeling capabilities. We combine this method with FFN, embedding dimension, and layer pruning, along with knowledge distillation-based retraining to obtain a unified compression recipe for hybrid models. Using this recipe, we compress the Nemotron-H 8B Hybrid model down to 4B parameters with up to 40× fewer training tokens compared to similarly-sized models. The resulting model surpasses the accuracy of similarly-sized models while achieving ∼ 2× faster inference throughput, significantly advancing the Pareto frontier.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Pluggable Pruning with Contiguous Layer Distillation for Diffusion TransformersJian Ma, Qirong Peng, Xujie Zhu, Peixing Xie 等CVPR 2026 · 被引用 7 次
- UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMsHung-Yueh Chiang, Chi-Chih Chang, Yu-Chen Lu, Chien-Yu Lin 等ICLR 2026 · 被引用 6 次
- SparseSSM: Efficient Selective Structured State Space Models Can Be Pruned in One-ShotKaiwen TUO, Huan WangICML 2026 · 被引用 6 次
- Hankel Singular Value Regularization for Highly Compressible State Space ModelsPaul Schwerdtner, Jules Berman, Benjamin PeherstorferNeurIPS 2025 · 被引用 3 次
- Star Elastic: Many-in-One Reasoning LLMs with Efficient Budget ControlAli Taghibakhshi, Ruisi Cai, Saurav Muralidharan, Sharath Turuvekere Sreenivas 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper6
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 被引用 1,407 次
- Sheared LLaMA: Accelerating Language Model Pre-training via Structured PruningMengzhou Xia, Tianyu Gao, Zhiyuan Zeng, Danqi ChenICLR 2024 · 被引用 453 次
- SliceGPT: Compress Large Language Models by Deleting Rows and ColumnsSaleh Ashkboos, Maximilian L. Croci, Marcelo Gennari Do Nascimento, Torsten Hoefler 等ICLR 2024 · 被引用 346 次
- Compact Language Models via Pruning and Knowledge DistillationSaurav Muralidharan, Sharath Turuvekere Sreenivas, Raviraj Joshi, Marcin Chochowski 等NeurIPS 2024 · 被引用 198 次
- Fluctuation-Based Adaptive Structured Pruning for Large Language ModelsYongqi An, Xu Zhao, Tao Yu, Ming Tang 等AAAI 2024 · 被引用 130 次
相关 Paper
- Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic ModelsAviv Bick, Kevin Y. Li, Eric P. Xing, J. Zico Kolter 等NeurIPS 2024 · 被引用 78 次
- MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language ModelingYingyue Li, Bencheng Liao, Wenyu Liu, Xinggang WangICCV 2025 · 被引用 1 次
- Zebra-Llama: Towards Extremely Efficient Hybrid ModelsMingyu Yang, Mehdi Rezagholizadeh, Guihong Li, Vikram Appia 等NeurIPS 2025 · 被引用 19 次
- Gradient-based Intra-attention Pruning on Pre-trained Language ModelsZiqing Yang, Yiming Cui, Xin Yao, Shijin WangACL 2023 · 被引用 2 次
- Efficient Unstructured Pruning of Mamba State-Space Models for Resource-Constrained EnvironmentsIbne Farabi Shihab, Sanjeda Akter, Anuj SharmaEMNLP 2025 · 被引用 1 次
