A Memory-Efficient Edge Inference Accelerator with XOR-based Model Compression
Hyunseung Lee, Jihoon Hong, Soosung Kim, Seung Yul Lee, Jae W. Lee
摘要
Model compression is widely adopted for edge inference of neural networks (NNs) to minimize both costly DRAM accesses and memory footprints. Recently, XOR-based model compression has demonstrated promising results to maximize compression ratio and minimize accuracy drop. However, XOR-based decompression alone produces bit errors and requires auxiliary data for error correction. To minimize model size and hence DRAM traffic, we propose an enhanced decompression algorithm and a low-cost hardware accelerator for it. Since not all errors are equal, our algorithm selects only important errors to correct with no accuracy drop. Compared with the baseline XOR compression scheme correcting all errors, the compressed model size of ResNet-18 and VGG-16 is reduced by 23% and 27% respectively. We also present a low-cost hardware implementation of on-line XOR decompression and error-correction logic built on Gemmini, an open-source systolic array accelerator, at the cost of only a 0.39% and 0.46% increase in area and power.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Gemmini: Enabling Systematic Deep-Learning Architecture Evaluation via Full-Stack IntegrationHasan Genc, Seah Kim, Alon Amid, Ameer Haj-Ali 等DAC 2021 · 被引用 325 次
- Sparse GPU kernels for deep learningTrevor Gale, Matei Zaharia, Cliff Young, Erich ElsenSC 2020 · 被引用 170 次
- SpZip: Architectural Support for Effective Data Compression In Irregular ApplicationsYifan Yang, Joel S. Emer, Daniel SánchezISCA 2021 · 被引用 31 次
- Encoding Weights of Irregular Sparsity for Fixed-to-Fixed Model CompressionBaeseong Park, Se Jung Kwon, Daehwan Oh, Byeongwook Kim 等ICLR 2022 · 被引用 4 次
- Structured Compression by Weight Encryption for Unstructured Pruning and QuantizationSe Jung Kwon, Dongsoo Lee, Byeongwook Kim, Parichay Kapoor 等CVPR 2020
相关 Paper
- MST-compression: Compressing and Accelerating Binary Neural Networks with Minimum Spanning TreeQuang Hieu Vo, Linh-Tam Tran, Sung-Ho Bae, Lok-Won Kim 等ICCV 2023 · 被引用 2 次
- Efficient Tunstall Decoder for Deep Neural Network CompressionChunyun Chen, Zhe Wang, Xiaowei Chen, Jie Lin 等DAC 2021 · 被引用 4 次
- Atalanta: A Bit is Worth a "Thousand" Tensor ValuesAlberto Delmas Lascorz, Mostafa Mahmoud, Ali Hadi Zadeh, Milos Nikolic 等ASPLOS 2024 · 被引用 7 次
- FleXOR: Trainable Fractional QuantizationDongsoo Lee, Se Jung Kwon, Byeongwook Kim, Yongkweon Jeon 等NeurIPS 2020 · 被引用 14 次
- An Efficient Deep Learning Accelerator for Compressed Video AnalysisYongchen Wang, Ying Wang, Huawei Li, Yinhe Han 等DAC 2020 · 被引用 4 次
