Compressing Deep Convolutional Neural Networks by Stacking Low-dimensional Binary Convolution Filters
Weichao Lan, Liang Lan
Abstract
Deep Convolutional Neural Networks (CNN) have been successfully applied to many real-life problems. However, the huge memory cost of deep CNN models poses a great challenge of deploying them on memory-constrained devices (e.g., mobile phones). One popular way to reduce the memory cost of deep CNN model is to train binary CNN where the weights in convolution filters are either 1 or -1 and therefore each weight can be efficiently stored using a single bit. However, the compression ratio of existing binary CNN models is upper bounded by ∼ 32. To address this limitation, we propose a novel method to compress deep CNN model by stacking low-dimensional binary convolution filters. Our proposed method approximates a standard convolution filter by selecting and stacking filters from a set of low-dimensional binary convolution filters. This set of low-dimensional binary convolution filters is shared across all filters for a given convolution layer. Therefore, our method will achieve much larger compression ratio than binary CNN models. In order to train our proposed model, we have theoretically shown that our proposed model is equivalent to select and stack intermediate feature maps generated by low-dimensional binary filters. Therefore, our proposed model can be efficiently trained using the split-transform-merge strategy. We also provide detailed analysis of the memory and computation cost of our model in model inference. We compared the proposed method with other five popular model compression techniques on two benchmark datasets. Our experimental results have demonstrated that our proposed method achieves much higher compression ratio than existing methods while maintains comparable accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- S2NN: Sub-bit Spiking Neural NetworksWenjie Wei, Malu Zhang, Jieyuan Zhang, Ammar Belatreche et al.NeurIPS 2025 · 1 citation
- Compacting Binary Neural Networks by Sparse Kernel SelectionYikai Wang, Wenbing Huang, Yinpeng Dong, Fuchun Sun et al.CVPR 2023
Related papers
- MST-compression: Compressing and Accelerating Binary Neural Networks with Minimum Spanning TreeQuang Hieu Vo, Linh-Tam Tran, Sung-Ho Bae, Lok-Won Kim et al.ICCV 2023 · 2 citations
- Sub-bit Neural Networks: Learning to Compress and Accelerate Binary Neural NetworksYikai Wang, Yi Yang, Fuchun Sun, Anbang YaoICCV 2021 · 18 citations
- DIVISION: Memory Efficient Training via Dual Activation PrecisionGuanchu Wang, Zirui Liu, Zhimeng Jiang, Ninghao Liu et al.ICML 2023 · 4 citations
- Adaptive Loss-Aware Quantization for Multi-Bit NetworksZhongnan Qu, Zimu Zhou, Yun Cheng, Lothar ThieleCVPR 2020
- FSNet: Compression of Deep Convolutional Neural Networks by Filter SummaryYingzhen Yang, Jiahui Yu, Nebojsa Jojic, Jun Huan et al.ICLR 2020 · 19 citations
