NxMTransformer: Semi-Structured Sparsification for Natural Language Understanding via ADMM
Connor Holmes, Minjia Zhang, Yuxiong He, Bo Wu
摘要
Natural Language Processing (NLP) has recently achieved great success by using huge pre-trained Transformer networks. However, these models often contain hundreds of millions or even billions of parameters, bringing challenges to online deployment due to latency constraints. Recently, hardware manufacturers have introduced dedicated hardware for NxM sparsity to provide the flexibility of unstructured pruning with the runtime efficiency of structured approaches. NxM sparsity permits arbitrarily selecting M parameters to retain from a contiguous group of N in the dense representation. However, due to the extremely high complexity of pre-trained models, the standard sparse fine-tuning techniques often fail to generalize well on downstream tasks, which have limited data resources. To address such an issue in a principled manner, we introduce a new learning framework, called NxMTransformer, to induce NxM semi-structured sparsity on pretrained language models for natural language understanding to obtain better performance. In particular, we propose to formulate the NxM sparsity as a constrained optimization problem and use Alternating Direction Method of Multipliers (ADMM) to optimize the downstream tasks while taking the underlying hardware constraints into consideration. ADMM decomposes the NxM sparsification problem into two sub-problems that can be solved sequentially, generating sparsified Transformer networks that achieve high accuracy while being able to effectively execute on newly released hardware. We apply our approach to a wide range of NLP tasks, and our proposed method is able to achieve 1.7 points higher accuracy in GLUE score than current best practices. Moreover, we perform detailed analysis on our approach and shed light on how ADMM affects fine-tuning accuracy for downstream tasks. Finally, we illustrate how NxMTransformer achieves additional performance improvement with knowledge distillation based methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- STEP: Learning N: M Structured Sparsity Masks from Scratch with PreconditionYucheng Lu, Shivani Agrawal, Suvinay Subramanian, Oleg Rybakov 等ICML 2023 · 被引用 32 次
- DoRA: Enhancing Parameter-Efficient Fine-Tuning with Dynamic Rank DistributionYulong Mao, Kaiyu Huang, Changhao Guan, Ganglin Bao 等ACL 2024 · 被引用 15 次
- Decouple then Classify: A Dynamic Multi-view Labeling Strategy with Shared and Specific InformationXinhang Wan, Jiyuan Liu, Xinwang Liu, Yi Wen 等ICML 2024 · 被引用 6 次
- Samoyeds: Accelerating MoE Models with Structured Sparsity Leveraging Sparse Tensor CoresChenpeng Wu, Qiqi Gu, Heng Shi, Jianguo Yao 等EuroSys 2025 · 被引用 5 次
- Privacy Amplification by Iteration for ADMM with (Strongly) Convex Objective FunctionsT.-H. Hubert Chan, Hao Xie, Mengshi ZhaoAAAI 2024 · 被引用 1 次
它引用的顶会 Paper7
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu 等ACL 2020 · 被引用 660 次
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERTSheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma 等AAAI 2020 · 被引用 656 次
- The Lottery Ticket Hypothesis for Pre-trained BERT NetworksTianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu 等NeurIPS 2020 · 被引用 428 次
- Learning N: M Fine-grained Structured Sparse Neural Networks From ScratchAojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu 等ICLR 2021 · 被引用 301 次
相关 Paper
- Pruning Large Language Models with Semi-Structural Adaptive Sparse TrainingWeiyu Huang, Yuezhou Hu, Guohao Jian, Jun Zhu 等AAAI 2025 · 被引用 25 次
- Learning Semi-Structured Sparsity for LLMs via Shared and Context-Aware HypernetworkLu Sun, Jun SakumaICLR 2026
- FNM-Trans: Efficient FPGA-based Transformer Architecture with Full N: M SparsityManting Zhang, Jialin Cao, Kejia Shi, Keqing Zhao 等DAC 2024 · 被引用 10 次
- MLX: Multi-Layer Execution for Structured LLM Workload Acceleration on Spatial ArchitecturesHaibin Wu, Wenming Li, Zhihua Fan, Zirui Ma 等ISCA 2026
- Gradient-based Intra-attention Pruning on Pre-trained Language ModelsZiqing Yang, Yiming Cui, Xin Yao, Shijin WangACL 2023 · 被引用 2 次
