Partially-Structured Transformer Pruning with Patch-Limited XOR-Gate Compression for Stall-Free Sparse-Model Access
Younghoon Byun, Youngjoo Lee
摘要
The pruning-based model compression is regarded as an essential technique to deploy the recent large-size transformer models in practical services; however, accessing sparse transformer models cannot reach the ideal speed at all due to the frequent memory stalls for the irregular memory-accessing patterns. Based on the recent XOR-gate compression relaxing the amount of irregular accesses, this work presents a novel partially-structured transformer pruning method dedicated to the interface-friendly compression format. The stall-free memory access is firstly derived by limiting the number of patches per weight, introducing a new trade-off between model quality and effective memory bandwidth. Then, the partially-structured pruning patterns are deployed to provide better accuracy-bandwidth trade-off by significantly reducing the number of correction patches. Adjusting the patch distribution per weight in an aggressive way, the number of limited patches can be even smaller than that of weight bits, further increasing the effective bandwidth for achieving the similar model accuracy. We demonstrate the proposed stall-free XOR-gate compression schemes at pruned DeiT/BERT models on ImageNet/SQuAD datasets, presenting the highest effective bandwidth for accessing sparse transformers compared to the existing stall-based solutions.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Structured Compression by Weight Encryption for Unstructured Pruning and QuantizationSe Jung Kwon, Dongsoo Lee, Byeongwook Kim, Parichay Kapoor 等CVPR 2020
- Deep Compression of Pre-trained Transformer ModelsNaigang Wang, Chi-Chun (Charlie) Liu, Swagath Venkataramani, Sanchari Sen 等NeurIPS 2022 · 被引用 38 次
- GOHSP: A Unified Framework of Graph and Optimization-Based Heterogeneous Structured Pruning for Vision TransformerMiao Yin, Burak Uzkent, Yilin Shen, Hongxia Jin 等AAAI 2023 · 被引用 23 次
- Coruscant: Co-Designing GPU Kernel and Sparse Tensor Core to Advocate Unstructured Sparsity in Efficient LLM InferenceDonghyeon Joo, Helya Hosseini, Ramyad Hadidi, Bahar AsgariMICRO 2025 · 被引用 8 次
- The Optimal BERT Surgeon: Scalable and Accurate Second-Order Pruning for Large Language ModelsEldar Kurtic, Daniel Campos, Tuan Nguyen, Elias Frantar 等EMNLP 2022 · 被引用 4 次
