Structured Compression by Weight Encryption for Unstructured Pruning and Quantization
Se Jung Kwon, Dongsoo Lee, Byeongwook Kim, Parichay Kapoor, Baeseong Park, Gu-Yeon Wei
摘要
Model compression techniques, such as pruning and quantization, are becoming increasingly important to reduce the memory footprints and the amount of computations. Despite model size reduction, achieving performance enhancement on devices is, however, still challenging mainly due to the irregular representations of sparse matrix formats. This paper proposes a new weight representation scheme for Sparse Quantized Neural Networks, specifically achieved by fine-grained and unstructured pruning method. The representation is encrypted in a structured regular format, which can be efficiently decoded through XOR-gate network during inference in a parallel manner. We demonstrate various deep learning models that can be compressed and represented by our proposed format with fixed and high compression ratio. For example, for fully-connected layers of AlexNet on ImageNet dataset, we can represent the sparse weights by only 0.28 bits/weight for 1-bit quantization and 91% pruning rate with a fixed decoding rate and full memory bandwidth usage. Decoding through XOR-gate network can be performed without any model accuracy degradation with additional patch data associated with small overhead.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Deep Compression of Pre-trained Transformer ModelsNaigang Wang, Chi-Chun (Charlie) Liu, Swagath Venkataramani, Sanchari Sen 等NeurIPS 2022 · 被引用 38 次
- FleXOR: Trainable Fractional QuantizationDongsoo Lee, Se Jung Kwon, Byeongwook Kim, Yongkweon Jeon 等NeurIPS 2020 · 被引用 14 次
- Enhancing Low-Rank Adaptation with Recoverability-Based Reinforcement Pruning for Object CountingHaojie Guo, Junyu Gao, Yuan YuanAAAI 2025 · 被引用 4 次
- Encoding Weights of Irregular Sparsity for Fixed-to-Fixed Model CompressionBaeseong Park, Se Jung Kwon, Daehwan Oh, Byeongwook Kim 等ICLR 2022 · 被引用 4 次
- A Memory-Efficient Edge Inference Accelerator with XOR-based Model CompressionHyunseung Lee, Jihoon Hong, Soosung Kim, Seung Yul Lee 等DAC 2023 · 被引用 4 次
相关 Paper
- Tight Compression: Compressing CNN Model Tightly Through Unstructured Pruning and Simulated Annealing Based PermutationXizi Chen, Jingyang Zhu, Jingbo Jiang, Chi-Ying TsuiDAC 2020 · 被引用 30 次
- Partially-Structured Transformer Pruning with Patch-Limited XOR-Gate Compression for Stall-Free Sparse-Model AccessYounghoon Byun, Youngjoo LeeDAC 2024
- Unified Data-Free Compression: Pruning and Quantization without Fine-TuningShipeng Bai, Jun Chen, Xintian Shen, Yixuan Qian 等ICCV 2023 · 被引用 31 次
- Harmonious Coexistence of Structured Weight Pruning and Ternarization for Deep Neural NetworksLi Yang, Zhezhi He, Deliang FanAAAI 2020 · 被引用 28 次
- BSQ: Exploring Bit-Level Sparsity for Mixed-Precision Neural Network QuantizationHuanrui Yang, Lin Duan, Yiran Chen, Hai LiICLR 2021 · 被引用 83 次
