ARMOR: High-Performance Semi-Structured Pruning via Adaptive Matrix Factorization
Lawrence Liu, Alexander Liu, Mengdi Wang, Tuo Zhao, Lin Yang
摘要
Large language models (LLMs) present significant deployment challenges due to their immense computational and memory requirements. While semi-structured pruning, particularly 2:4 sparsity, offers a path to practical hardware acceleration, existing methods often incur substantial performance degradation. To bridge this gap, we introduce ARMOR: (Adaptive Representation with Matrix- factORization), a novel one-shot post-training pruning algorithm. Instead of directly pruning weights, ARMOR factorizes each weight matrix into a 2:4 sparse core wrapped by two low-overhead, block diagonal matrices. These wrappers act as efficient pre- and post-transformation error correctors, offering greater flexibility to preserve model quality compared to conventional 2:4 pruning techniques. The sparse core and block diagonal wrappers are chosen through a block coordinate descent algorithm that minimizes a layer-wise proxy loss. We prove this optimization is guaranteed to converge to a solution with a proxy loss less than or equal to state-of-the-art pruning algorithms. Experiments on Llama (Touvron et al., 2023; Dubey et al., 2024) and Qwen (Yang et al., 2025) model families demonstrate that ARMOR consistently and significantly outperforms state-of-the-art 2:4 pruning methods across a wide range of downstream tasks and perplexity evaluations, and generalizes to provide improvements for general N:M patterns and unstructured sparsity. ARMOR achieves this superior performance while retaining the inference speedups and substantial memory usage reductions of 2:4 pruning, establishing a more effective trade-off between model compression and task accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Unified Static-Dynamic Pruning for Efficient LLM InferenceJinhyeok Kim, Yejoon Lee, Jaeyoung DoVLDB 2026
- Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via SpontaneityHaotian Xu, Jiannan Yang, Tian Gao, Lily Weng 等ICML 2026
它引用的顶会 Paper17
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 被引用 1,240 次
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 被引用 994 次
相关 Paper
- Learning Semi-Structured Sparsity for LLMs via Shared and Context-Aware HypernetworkLu Sun, Jun SakumaICLR 2026
- FISTAPruner: Layer-wise Post-training Pruning for Large Language ModelsPengxiang Zhao, Hanyu Hu, Ping Li, Yi Zheng 等EMNLP 2025
- WRP: Weight Recover Prune for Structured SparsityZhendong Tan, Xingjun Zhang, Zheng WeiACL 2024
- SlideSparse: Fast and Flexible (2N-2):2N Structured SparsityYingbo HAO, Hanyong Shao, Ting Song, Yan Xia 等ICML 2026
- DELTA4: Sparse Matrix-Vector Multiplication for Low SparsityVladimír Macko, Vladimír BožaICML 2026 · 被引用 9 次
