Understanding weight-magnitude hyperparameters in training binary networks
Joris Quist, Yunqiang Li, Jan van Gemert
Abstract
Binary Neural Networks (BNNs) are compact and efficient by using binary weights instead of real-valued weights. Current BNNs use latent real-valued weights during training, where several training hyper-parameters are inherited from real-valued networks. The interpretation of several of these hyperparameters is based on the magnitude of the real-valued weights. For BNNs, however, the magnitude of binary weights is not meaningful, and thus it is unclear what these hyperparameters actually do. One example is weight-decay, which aims to keep the magnitude of real-valued weights small. Other examples are latent weight initialization, the learning rate, and learning rate decay, which influence the magnitude of the real-valued weights. The magnitude is interpretable for real-valued weights, but loses its meaning for binary weights. In this paper we offer a new interpretation of these magnitude-based hyperparameters based on higher-order gradient filtering during network optimization. Our analysis makes it possible to understand how magnitude-based hyperparameters influence the training of binary networks which allows for new optimization filters specifically designed for binary neural networks that are independent of their real-valued interpretation. Moreover, our improved understanding reduces the number of hyperparameters, which in turn eases the hyperparameter tuning effort which may lead to better hyperparameter values for improved accuracy. Code is available at https://github.com/jorisquist/Understanding-WM-HP-in-BNNs
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on14
- BoTorch: A Framework for Efficient Monte-Carlo Bayesian OptimizationMaximilian Balandat, Brian Karrer, Daniel R. Jiang, Samuel Daulton et al.NeurIPS 2020 · 686 citations
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural NetworksRuihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li et al.ICCV 2019 · 540 citations
- Training binary neural networks with real-to-binary convolutionsBrais Martínez, Jing Yang, Adrian Bulat, Georgios TzimiropoulosICLR 2020 · 251 citations
- ReCU: Reviving the Dead Weights in Binary Neural NetworksZihan Xu, Mingbao Lin, Jianzhuang Liu, Jie Chen et al.ICCV 2021 · 102 citations
- How Do Adam and Training Strategies Help BNNs OptimizationZechun Liu, Zhiqiang Shen, Shichao Li, Koen Helwegen et al.ICML 2021 · 100 citations
Related papers
- Resilient Binary Neural NetworkSheng Xu, Yanjing Li, Teli Ma, Mingbao Lin et al.AAAI 2023 · 1 citation
- Reducing the Computational Cost of Deep Generative Models with Binary Neural NetworksThomas Bird, Friso H. Kingma, David BarberICLR 2021 · 3 citations
- Sub-bit Neural Networks: Learning to Compress and Accelerate Binary Neural NetworksYikai Wang, Yi Yang, Fuchun Sun, Anbang YaoICCV 2021 · 18 citations
- Training Binary Neural Networks via Gaussian Variational Inference and Low-Rank Semidefinite ProgrammingLorenzo Orecchia, Jiawei Hu, Xue He, Wang Mark et al.NeurIPS 2024 · 4 citations
- BEP: A Binary Error Propagation Algorithm for Binary Neural Networks TrainingLuca Colombo, Fabrizio Pittorino, Daniele Zambon, Carlo Baldassi et al.ICLR 2026
