Deterministic Differentiable Structured Pruning for Large Language Models
Weiyu Huang, Pengle Zhang, Xiaolu Zhang, JUN ZHOU, Jun Zhu, Jianfei Chen
摘要
Structured pruning reduces LLM inference cost by removing low-importance architectural components. This can be viewed as learning a multiplicative gate for each component under an ℓ 0 sparsity constraint. Due to the discreteness of the ℓ 0 norm, prior work typically adopts stochastic hard-concrete relaxations to enable differentiable optimization; however, this stochasticity can introduce a train-test mismatch when sampled masks are discretized for deployment and restricts masks to a bounded, near-binary range. To address this, we propose Deterministic Differentiable Pruning (DDP), a mask-only optimization method that eliminates stochasticity by directly optimizing a deterministic soft surrogate of the discrete ℓ 0 objective. Compared with prior approaches, DDP offers greater expressiveness, reduced train-test mismatch, and faster convergence. We apply our method to several dense and MoE models, including Qwen3-32B and Qwen3-30B-A3B, achieving a performance loss as small as 1% on downstream tasks while outperforming previous methods at 20% sparsity. We further demonstrate end-to-end inference speedups in realistic deployment settings with vLLM. Our code is publicly available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 被引用 794 次
- Structured Pruning of Large Language ModelsZiheng Wang, Jeremy Wohlwend, Tao LeiEMNLP 2020 · 被引用 88 次
相关 Paper
- Bypass Back-propagation: Optimization-based Structural Pruning for Large Language Models via Policy GradientYuan Gao, Zujing Liu, Weizhong Zhang, Bo Du 等ACL 2025
- From Local to Global: Revisiting Structured Pruning Paradigms for Large Language ModelsZiyan Wang, Enmao Diao, Qi Le, Pu Wang 等ACL 2026 · 被引用 2 次
- Learning Semi-Structured Sparsity for LLMs via Shared and Context-Aware HypernetworkLu Sun, Jun SakumaICLR 2026
- Discrete Model Compression With Resource Constraint for Deep Neural NetworksShangqian Gao, Feihu Huang, Jian Pei, Heng HuangCVPR 2020
- Dynamic Sparse No Training: Training-Free Fine-tuning for Sparse LLMsYuxin Zhang, Lirui Zhao, Mingbao Lin, Yunyun Sun 等ICLR 2024 · 被引用 78 次
