ReCU: Reviving the Dead Weights in Binary Neural Networks
Zihan Xu, Mingbao Lin, Jianzhuang Liu, Jie Chen, Ling Shao, Yue Gao, Yonghong Tian, Rongrong Ji
Abstract
Binary neural networks (BNNs) have received increasing attention due to their superior reductions of computation and memory. Most existing works focus on either lessening the quantization error by minimizing the gap between the full-precision weights and their binarization or designing a gradient approximation to mitigate the gradient mismatch, while leaving the "dead weights" untouched. This leads to slow convergence when training BNNs. In this paper, for the first time, we explore the influence of "dead weights" which refer to a group of weights that are barely updated during the training of BNNs, and then introduce rectified clamp unit (ReCU) to revive the "dead weights" for updating. We prove that reviving the "dead weights" by ReCU can result in a smaller quantization error. Besides, we also take into account the information entropy of the weights, and then mathematically analyze why the weight standardization can benefit BNNs. We demonstrate the inherent contradiction between minimizing the quantization error and maximizing the information entropy, and then propose an adaptive exponential scheduler to identify the range of the "dead weights". By considering the "dead weights", our method offers not only faster BNN training, but also state-of-the-art performance on CIFAR-10 and Im-ageNet, compared with recent methods. Code can be available at https://github.com/z-hXu/ReCU .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6495355e-da05-4f8f-b2f1-59e013ef52b7Cited by top-tier papers17
- BiBERT: Accurate Fully Binarized BERTHaotong Qin, Yifu Ding, Mingyuan Zhang, Qinghua Yan et al.ICLR 2022 · 121 citations
- PB-LLM: Partially Binarized Large Language ModelsZhihang Yuan, Yuzhang Shang, Zhen DongICLR 2024 · 91 citations
- BiMatting: Efficient Video Matting via BinarizationHaotong Qin, Lei Ke, Xudong Ma, Martin Danelljan et al.NeurIPS 2023 · 28 citations
- INSTA-BNN: Binary Neural Network with INSTAnce-aware ThresholdChanghun Lee, Hyungjun Kim, Eunhyeok Park, Jae-Joon KimICCV 2023 · 16 citations
- BiDM: Pushing the Limit of Quantization for Diffusion ModelsXingyu Zheng, Xianglong Liu, Yichen Bian, Xudong Ma et al.NeurIPS 2024 · 12 citations
Builds on9
- MetaPruning: Meta Learning for Automatic Neural Network Channel PruningZechun Liu, Haoyuan Mu, Xiangyu Zhang, Zichao Guo et al.ICCV 2019 · 633 citations
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural NetworksRuihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li et al.ICCV 2019 · 540 citations
- Rotated Binary Neural NetworkMingbao Lin, Rongrong Ji, Zihan Xu, Baochang Zhang et al.NeurIPS 2020 · 161 citations
- Searching for Low-Bit Weights in Quantized Neural NetworksZhaohui Yang, Yunhe Wang, Kai Han, Chunjing Xu et al.NeurIPS 2020 · 103 citations
- Bayesian Optimized 1-Bit CNNsJiaxin Gu, Junhe Zhao, Xiaolong Jiang, Baochang Zhang et al.ICCV 2019 · 57 citations
Related papers
- SA-BNN: State-Aware Binary Neural NetworkChunlei Liu, Peng Chen, Bohan Zhuang, Chunhua Shen et al.AAAI 2021 · 23 citations
- BEP: A Binary Error Propagation Algorithm for Binary Neural Networks TrainingLuca Colombo, Fabrizio Pittorino, Daniele Zambon, Carlo Baldassi et al.ICLR 2026
- Estimator Meets Equilibrium Perspective: A Rectified Straight Through Estimator for Binary Neural Networks TrainingXiao-Ming Wu, Dian Zheng, Zuhao Liu, Wei-Shi ZhengICCV 2023 · 28 citations
- BinaryDuo: Reducing Gradient Mismatch in Binary Activation Network by Coupling Binary ActivationsHyungjun Kim, Kyungsu Kim, Jinseok Kim, Jae-Joon KimICLR 2020 · 51 citations
- Resilient Binary Neural NetworkSheng Xu, Yanjing Li, Teli Ma, Mingbao Lin et al.AAAI 2023 · 1 citation
