Sub-bit Neural Networks: Learning to Compress and Accelerate Binary Neural Networks
Yikai Wang, Yi Yang, Fuchun Sun, Anbang Yao
摘要
In the low-bit quantization field, training Binarized Neural Networks (BNNs) is the extreme solution to ease the deployment of deep models on resource-constrained devices, having the lowest storage cost and significantly cheaper bit-wise operations compared to 32-bit floating-point counterparts. In this paper, we introduce Sub-bit Neural Networks (SNNs), a new type of binary quantization design tailored to compress and accelerate BNNs. SNNs are inspired by an empirical observation, showing that binary kernels learnt at convolutional layers of a BNN model are likely to be distributed over kernel subsets. As a result, unlike existing methods that binarize weights one by one, SNNs are trained with a kernel-aware optimization framework, which exploits binary quantization in the fine-grained convolutional kernel space. Specifically, our method includes a random sampling step generating layer-specific subsets of the kernel space, and a refinement step learning to adjust these subsets of binary kernels via optimization. Experiments on visual recognition benchmarks and the hardware deployment on FPGA validate the great potentials of SNNs. For instance, on ImageNet, SNNs of ResNet-18/ResNet-34 with 0.56-bit weights achieve 3.13/3.33× runtime speedup and 1.8× compression over conventional BNNs with moderate drops in recognition accuracy. Promising results are also obtained when applying SNNs to binarize both weights and activations. Our code is available at https://github.com/yikaiw/SNN.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- SmartLite: A DBMS-based Serving System for DNN Inference in Resource-constrained EnvironmentsQiuru Lin, Sai Wu, Junbo Zhao, Jian Dai 等VLDB 2024 · 被引用 17 次
- MST-compression: Compressing and Accelerating Binary Neural Networks with Minimum Spanning TreeQuang Hieu Vo, Linh-Tam Tran, Sung-Ho Bae, Lok-Won Kim 等ICCV 2023 · 被引用 2 次
- S2NN: Sub-bit Spiking Neural NetworksWenjie Wei, Malu Zhang, Jieyuan Zhang, Ammar Belatreche 等NeurIPS 2025 · 被引用 1 次
- STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMsPeijie Dong, Lujun Li, Yuedong Zhong, Dayou Du 等ICLR 2025 · 被引用 1 次
- Compacting Binary Neural Networks by Sparse Kernel SelectionYikai Wang, Wenbing Huang, Yinpeng Dong, Fuchun Sun 等CVPR 2023
它引用的顶会 Paper6
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
- Universally Slimmable Networks and Improved Training TechniquesJiahui Yu, Thomas S. HuangICCV 2019 · 被引用 444 次
- Rotated Binary Neural NetworkMingbao Lin, Rongrong Ji, Zihan Xu, Baochang Zhang 等NeurIPS 2020 · 被引用 161 次
- FleXOR: Trainable Fractional QuantizationDongsoo Lee, Se Jung Kwon, Byeongwook Kim, Yongkweon Jeon 等NeurIPS 2020 · 被引用 14 次
- BiDet: An Efficient Binarized Object DetectorZiwei Wang, Ziyi Wu, Jiwen Lu, Jie ZhouCVPR 2020
相关 Paper
- Adaptive Loss-Aware Quantization for Multi-Bit NetworksZhongnan Qu, Zimu Zhou, Yun Cheng, Lothar ThieleCVPR 2020
- RTN: Reparameterized Ternary NetworkYuhang Li, Xin Dong, Sai Qian Zhang, Haoli Bai 等AAAI 2020 · 被引用 34 次
- Fast and Accurate Binary Neural Networks Based on Depth-Width ReshapingPing Xue, Yang Lu, Jingfei Chang, Xing Wei 等AAAI 2023 · 被引用 3 次
- Improving Low-Precision Network Quantization via Bin RegularizationTiantian Han, Dong Li, Ji Liu, Lu Tian 等ICCV 2021 · 被引用 45 次
- Training Binary Neural Network without Batch Normalization for Image Super-ResolutionXinrui Jiang, Nannan Wang, Jingwei Xin, Keyu Li 等AAAI 2021 · 被引用 52 次
