Structured Compression by Weight Encryption for Unstructured Pruning and Quantization
Se Jung Kwon, Dongsoo Lee, Byeongwook Kim, Parichay Kapoor, Baeseong Park, Gu-Yeon Wei
Abstract
Model compression techniques, such as pruning and quantization, are becoming increasingly important to reduce the memory footprints and the amount of computations. Despite model size reduction, achieving performance enhancement on devices is, however, still challenging mainly due to the irregular representations of sparse matrix formats. This paper proposes a new weight representation scheme for Sparse Quantized Neural Networks, specifically achieved by fine-grained and unstructured pruning method. The representation is encrypted in a structured regular format, which can be efficiently decoded through XOR-gate network during inference in a parallel manner. We demonstrate various deep learning models that can be compressed and represented by our proposed format with fixed and high compression ratio. For example, for fully-connected layers of AlexNet on ImageNet dataset, we can represent the sparse weights by only 0.28 bits/weight for 1-bit quantization and 91% pruning rate with a fixed decoding rate and full memory bandwidth usage. Decoding through XOR-gate network can be performed without any model accuracy degradation with additional patch data associated with small overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c89f500e-2d41-4ac2-8333-e175f42be1e4Cited by top-tier papers6
- Deep Compression of Pre-trained Transformer ModelsNaigang Wang, Chi-Chun (Charlie) Liu, Swagath Venkataramani, Sanchari Sen et al.NeurIPS 2022 · 38 citations
- FleXOR: Trainable Fractional QuantizationDongsoo Lee, Se Jung Kwon, Byeongwook Kim, Yongkweon Jeon et al.NeurIPS 2020 · 14 citations
- Enhancing Low-Rank Adaptation with Recoverability-Based Reinforcement Pruning for Object CountingHaojie Guo, Junyu Gao, Yuan YuanAAAI 2025 · 4 citations
- Encoding Weights of Irregular Sparsity for Fixed-to-Fixed Model CompressionBaeseong Park, Se Jung Kwon, Daehwan Oh, Byeongwook Kim et al.ICLR 2022 · 4 citations
- A Memory-Efficient Edge Inference Accelerator with XOR-based Model CompressionHyunseung Lee, Jihoon Hong, Soosung Kim, Seung Yul Lee et al.DAC 2023 · 4 citations
Related papers
- Tight Compression: Compressing CNN Model Tightly Through Unstructured Pruning and Simulated Annealing Based PermutationXizi Chen, Jingyang Zhu, Jingbo Jiang, Chi-Ying TsuiDAC 2020 · 30 citations
- Partially-Structured Transformer Pruning with Patch-Limited XOR-Gate Compression for Stall-Free Sparse-Model AccessYounghoon Byun, Youngjoo LeeDAC 2024
- Unified Data-Free Compression: Pruning and Quantization without Fine-TuningShipeng Bai, Jun Chen, Xintian Shen, Yixuan Qian et al.ICCV 2023 · 31 citations
- Harmonious Coexistence of Structured Weight Pruning and Ternarization for Deep Neural NetworksLi Yang, Zhezhi He, Deliang FanAAAI 2020 · 28 citations
- BSQ: Exploring Bit-Level Sparsity for Mixed-Precision Neural Network QuantizationHuanrui Yang, Lin Duan, Yiran Chen, Hai LiICLR 2021 · 83 citations
