ReCU: Reviving the Dead Weights in Binary Neural Networks
Zihan Xu, Mingbao Lin, Jianzhuang Liu, Jie Chen, Ling Shao, Yue Gao, Yonghong Tian, Rongrong Ji
摘要
Binary neural networks (BNNs) have received increasing attention due to their superior reductions of computation and memory. Most existing works focus on either lessening the quantization error by minimizing the gap between the full-precision weights and their binarization or designing a gradient approximation to mitigate the gradient mismatch, while leaving the "dead weights" untouched. This leads to slow convergence when training BNNs. In this paper, for the first time, we explore the influence of "dead weights" which refer to a group of weights that are barely updated during the training of BNNs, and then introduce rectified clamp unit (ReCU) to revive the "dead weights" for updating. We prove that reviving the "dead weights" by ReCU can result in a smaller quantization error. Besides, we also take into account the information entropy of the weights, and then mathematically analyze why the weight standardization can benefit BNNs. We demonstrate the inherent contradiction between minimizing the quantization error and maximizing the information entropy, and then propose an adaptive exponential scheduler to identify the range of the "dead weights". By considering the "dead weights", our method offers not only faster BNN training, but also state-of-the-art performance on CIFAR-10 and Im-ageNet, compared with recent methods. Code can be available at https://github.com/z-hXu/ReCU .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- BiBERT: Accurate Fully Binarized BERTHaotong Qin, Yifu Ding, Mingyuan Zhang, Qinghua Yan 等ICLR 2022 · 被引用 121 次
- PB-LLM: Partially Binarized Large Language ModelsZhihang Yuan, Yuzhang Shang, Zhen DongICLR 2024 · 被引用 91 次
- BiMatting: Efficient Video Matting via BinarizationHaotong Qin, Lei Ke, Xudong Ma, Martin Danelljan 等NeurIPS 2023 · 被引用 28 次
- INSTA-BNN: Binary Neural Network with INSTAnce-aware ThresholdChanghun Lee, Hyungjun Kim, Eunhyeok Park, Jae-Joon KimICCV 2023 · 被引用 16 次
- BiDM: Pushing the Limit of Quantization for Diffusion ModelsXingyu Zheng, Xianglong Liu, Yichen Bian, Xudong Ma 等NeurIPS 2024 · 被引用 12 次
它引用的顶会 Paper9
- MetaPruning: Meta Learning for Automatic Neural Network Channel PruningZechun Liu, Haoyuan Mu, Xiangyu Zhang, Zichao Guo 等ICCV 2019 · 被引用 633 次
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural NetworksRuihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li 等ICCV 2019 · 被引用 540 次
- Rotated Binary Neural NetworkMingbao Lin, Rongrong Ji, Zihan Xu, Baochang Zhang 等NeurIPS 2020 · 被引用 161 次
- Searching for Low-Bit Weights in Quantized Neural NetworksZhaohui Yang, Yunhe Wang, Kai Han, Chunjing Xu 等NeurIPS 2020 · 被引用 103 次
- Bayesian Optimized 1-Bit CNNsJiaxin Gu, Junhe Zhao, Xiaolong Jiang, Baochang Zhang 等ICCV 2019 · 被引用 57 次
相关 Paper
- SA-BNN: State-Aware Binary Neural NetworkChunlei Liu, Peng Chen, Bohan Zhuang, Chunhua Shen 等AAAI 2021 · 被引用 23 次
- BEP: A Binary Error Propagation Algorithm for Binary Neural Networks TrainingLuca Colombo, Fabrizio Pittorino, Daniele Zambon, Carlo Baldassi 等ICLR 2026
- Estimator Meets Equilibrium Perspective: A Rectified Straight Through Estimator for Binary Neural Networks TrainingXiao-Ming Wu, Dian Zheng, Zuhao Liu, Wei-Shi ZhengICCV 2023 · 被引用 28 次
- BinaryDuo: Reducing Gradient Mismatch in Binary Activation Network by Coupling Binary ActivationsHyungjun Kim, Kyungsu Kim, Jinseok Kim, Jae-Joon KimICLR 2020 · 被引用 51 次
- Resilient Binary Neural NetworkSheng Xu, Yanjing Li, Teli Ma, Mingbao Lin 等AAAI 2023 · 被引用 1 次
