Efficient Tunstall Decoder for Deep Neural Network Compression
Chunyun Chen, Zhe Wang, Xiaowei Chen, Jie Lin, Mohamed M. Sabry Aly
Abstract
Power and area-efficient deep neural network (DNN) designs are key in edge applications. Compact DNNs, via compression or quantization, enable such designs by significantly reducing memory footprint. Lossless entropy coding can further reduce the size of networks. It is then critical to provide hardware support for such entropy coding module to fully benefit from the resulted reduced memory requirement. In this work, we introduce Tunstall coding to compress the quantized weights. Tunstall coding can achieve high compression ratio as well as very fast decoding speed on various deep networks. We present two hardware-accelerated decoding techniques that provide streamlined decoding capabilities. We synthesize these designs targeting on FPGA. Results show that we achieve up to 6× faster decoding time versus state-of-the-art decoding methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Atalanta: A Bit is Worth a "Thousand" Tensor ValuesAlberto Delmas Lascorz, Mostafa Mahmoud, Ali Hadi Zadeh, Milos Nikolic et al.ASPLOS 2024 · 7 citations
- MST-compression: Compressing and Accelerating Binary Neural Networks with Minimum Spanning TreeQuang Hieu Vo, Linh-Tam Tran, Sung-Ho Bae, Lok-Won Kim et al.ICCV 2023 · 2 citations
- Rethinking Pruning for Accelerating Deep Inference At the EdgeDawei Gao, Xiaoxi He, Zimu Zhou, Yongxin Tong et al.KDD 2020 · 24 citations
- Mix and Match: A Novel FPGA-Centric Deep Neural Network Quantization FrameworkSung-En Chang, Yanyu Li, Mengshu Sun, Runbin Shi et al.HPCA 2021 · 125 citations
- Differentiable Weightless Neural NetworksAlan Tendler Leibel Bacellar, Zachary Susskind, Maurício Breternitz Jr., Eugene John et al.ICML 2024 · 34 citations
