Sub-bit Neural Networks: Learning to Compress and Accelerate Binary Neural Networks
Yikai Wang, Yi Yang, Fuchun Sun, Anbang Yao
Abstract
In the low-bit quantization field, training Binarized Neural Networks (BNNs) is the extreme solution to ease the deployment of deep models on resource-constrained devices, having the lowest storage cost and significantly cheaper bit-wise operations compared to 32-bit floating-point counterparts. In this paper, we introduce Sub-bit Neural Networks (SNNs), a new type of binary quantization design tailored to compress and accelerate BNNs. SNNs are inspired by an empirical observation, showing that binary kernels learnt at convolutional layers of a BNN model are likely to be distributed over kernel subsets. As a result, unlike existing methods that binarize weights one by one, SNNs are trained with a kernel-aware optimization framework, which exploits binary quantization in the fine-grained convolutional kernel space. Specifically, our method includes a random sampling step generating layer-specific subsets of the kernel space, and a refinement step learning to adjust these subsets of binary kernels via optimization. Experiments on visual recognition benchmarks and the hardware deployment on FPGA validate the great potentials of SNNs. For instance, on ImageNet, SNNs of ResNet-18/ResNet-34 with 0.56-bit weights achieve 3.13/3.33× runtime speedup and 1.8× compression over conventional BNNs with moderate drops in recognition accuracy. Promising results are also obtained when applying SNNs to binarize both weights and activations. Our code is available at https://github.com/yikaiw/SNN.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3414a33d-ead4-417a-b0b4-ded34f3ab3aeCited by top-tier papers6
- SmartLite: A DBMS-based Serving System for DNN Inference in Resource-constrained EnvironmentsQiuru Lin, Sai Wu, Junbo Zhao, Jian Dai et al.VLDB 2024 · 17 citations
- MST-compression: Compressing and Accelerating Binary Neural Networks with Minimum Spanning TreeQuang Hieu Vo, Linh-Tam Tran, Sung-Ho Bae, Lok-Won Kim et al.ICCV 2023 · 2 citations
- S2NN: Sub-bit Spiking Neural NetworksWenjie Wei, Malu Zhang, Jieyuan Zhang, Ammar Belatreche et al.NeurIPS 2025 · 1 citation
- STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMsPeijie Dong, Lujun Li, Yuedong Zhong, Dayou Du et al.ICLR 2025 · 1 citation
- Compacting Binary Neural Networks by Sparse Kernel SelectionYikai Wang, Wenbing Huang, Yinpeng Dong, Fuchun Sun et al.CVPR 2023
Builds on6
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- Universally Slimmable Networks and Improved Training TechniquesJiahui Yu, Thomas S. HuangICCV 2019 · 444 citations
- Rotated Binary Neural NetworkMingbao Lin, Rongrong Ji, Zihan Xu, Baochang Zhang et al.NeurIPS 2020 · 161 citations
- FleXOR: Trainable Fractional QuantizationDongsoo Lee, Se Jung Kwon, Byeongwook Kim, Yongkweon Jeon et al.NeurIPS 2020 · 14 citations
- BiDet: An Efficient Binarized Object DetectorZiwei Wang, Ziyi Wu, Jiwen Lu, Jie ZhouCVPR 2020
Related papers
- Adaptive Loss-Aware Quantization for Multi-Bit NetworksZhongnan Qu, Zimu Zhou, Yun Cheng, Lothar ThieleCVPR 2020
- RTN: Reparameterized Ternary NetworkYuhang Li, Xin Dong, Sai Qian Zhang, Haoli Bai et al.AAAI 2020 · 34 citations
- Fast and Accurate Binary Neural Networks Based on Depth-Width ReshapingPing Xue, Yang Lu, Jingfei Chang, Xing Wei et al.AAAI 2023 · 3 citations
- Improving Low-Precision Network Quantization via Bin RegularizationTiantian Han, Dong Li, Ji Liu, Lu Tian et al.ICCV 2021 · 45 citations
- Training Binary Neural Network without Batch Normalization for Image Super-ResolutionXinrui Jiang, Nannan Wang, Jingwei Xin, Keyu Li et al.AAAI 2021 · 52 citations
