Encoding Weights of Irregular Sparsity for Fixed-to-Fixed Model Compression
Baeseong Park, Se Jung Kwon, Daehwan Oh, Byeongwook Kim, Dongsoo Lee
Abstract
Even though fine-grained pruning techniques achieve a high compression ratio, conventional sparsity representations (such as CSR) associated with irregular sparsity degrade parallelism significantly. Practical pruning methods, thus, usually lower pruning rates (by structured pruning) to improve parallelism. In this paper, we study fixed-to-fixed (lossless) encoding architecture/algorithm to support fine-grained pruning methods such that sparse neural networks can be stored in a highly regular structure. We first estimate the maximum compression ratio of encoding-based compression using entropy. Then, as an effort to push the compression ratio to the theoretical maximum (by entropy), we propose a sequential fixed-to-fixed encoding scheme. We demonstrate that our proposed compression scheme achieves almost the maximum compression ratio for the Transformer and ResNet-50 pruned by various fine-grained pruning methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3480ea18-5f9b-4a6f-99fc-cf5dbf301a38Cited by top-tier papers2
- A Memory-Efficient Edge Inference Accelerator with XOR-based Model CompressionHyunseung Lee, Jihoon Hong, Soosung Kim, Seung Yul Lee et al.DAC 2023 · 4 citations
- Underspecification in Language Modeling Tasks: A Causality-Informed Study of Gendered Pronoun ResolutionEmily McMilinAAAI 2024 · 1 citation
Builds on5
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 743 citations
- Comparing Rewinding and Fine-tuning in Neural Network PruningAlex Renda, Jonathan Frankle, Michael CarbinICLR 2020 · 437 citations
- Learning N: M Fine-grained Structured Sparse Neural Networks From ScratchAojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu et al.ICLR 2021 · 301 citations
- Sparse GPU kernels for deep learningTrevor Gale, Matei Zaharia, Cliff Young, Erich ElsenSC 2020 · 170 citations
- Structured Compression by Weight Encryption for Unstructured Pruning and QuantizationSe Jung Kwon, Dongsoo Lee, Byeongwook Kim, Parichay Kapoor et al.CVPR 2020
Related papers
- PCNN: Pattern-based Fine-Grained Regular Pruning Towards Optimizing CNN AcceleratorsZhanhong Tan, Jiebo Song, Xiaolong Ma, Sia Huat Tan et al.DAC 2020 · 28 citations
- Tight Compression: Compressing CNN Model Tightly Through Unstructured Pruning and Simulated Annealing Based PermutationXizi Chen, Jingyang Zhu, Jingbo Jiang, Chi-Ying TsuiDAC 2020 · 30 citations
- DARB: A Density-Adaptive Regular-Block Pruning for Deep Neural NetworksAo Ren, Tao Zhang, Yuhao Wang, Sheng Lin et al.AAAI 2020 · 11 citations
- Unified Data-Free Compression: Pruning and Quantization without Fine-TuningShipeng Bai, Jun Chen, Xintian Shen, Yixuan Qian et al.ICCV 2023 · 31 citations
- Layerwise Sparse Coding for Pruned Deep Neural Networks with Extreme Compression RatioXiao Liu, Wenbin Li, Jing Huo, Lili Yao et al.AAAI 2020 · 13 citations
