Learning Semi-Structured Sparsity for LLMs via Shared and Context-Aware Hypernetwork
Lu Sun, Jun Sakuma
摘要
Large Language Models (LLMs) achieve state-of-the-art performance but are costly to deploy in resource-constrained environments. Pruning with semi-structured sparsity reduces computation and enables hardware acceleration, yet existing methods face a trade-off: one-shot approaches are efficient but heuristic, while optimization-based methods are accurate but expensive.
We introduce HyperPrune, a resource-efficient framework that directly optimizes sparsity. A lightweight hypernetwork, shared across layers and conditioned on learnable embeddings, generates structured masks in a one-shot, layer-wise manner. Continual pruning preserves cross-layer knowledge, and feature outlier regularization retains critical activations, unifying the strengths of heuristic and optimization-based methods.
Experiments on LLaMA-7B to 70B show state-of-the-art accuracy–sparsity trade-offs on a single A100 GPU, achieving higher efficiency, accuracy, and scalability than prior approaches. HyperPrune offers a practical, scalable, and hardware-friendly solution for structured LLM pruning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 被引用 1,240 次
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 被引用 794 次
相关 Paper
- Proxsparse: Regularized Learning of Semi-Structured Sparsity masks for Pretrained LLMSHongyi Liu, Rajarshi Saha, Zhen Jia, Youngsuk Park 等ICML 2025
- Lua-LLM: Learning Unstructured-Sparsity Allocation for Large Language ModelsMingge Lu, Jingwei Sun, Junqing Lin, Zechun Zhou 等NeurIPS 2025 · 被引用 1 次
- MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMsYan Sun, Qixin Zhang, Zhiyuan Yu, Xikun Zhang 等ICLR 2026 · 被引用 1 次
- FISTAPruner: Layer-wise Post-training Pruning for Large Language ModelsPengxiang Zhao, Hanyu Hu, Ping Li, Yi Zheng 等EMNLP 2025
- Two-Stage Regularization-Based Structured Pruning for LLMsMingkuan Feng, Jinyang Wu, Siyuan Liu, Shuai Zhang 等ACL 2026 · 被引用 3 次
