How Do Adam and Training Strategies Help BNNs Optimization
Zechun Liu, Zhiqiang Shen, Shichao Li, Koen Helwegen, Dong Huang, Kwang-Ting Cheng
Abstract
The best performing Binary Neural Networks (BNNs) are usually attained using Adam optimization and its multi-step training variants (Rastegari et al., 2016; Liu et al., 2020) . However, to the best of our knowledge, few studies explore the fundamental reasons why Adam is superior to other optimizers like SGD for BNN optimization or provide analytical explanations that support specific training strategies. To address this, in this paper we first investigate the trajectories of gradients and weights in BNNs during the training process. We show the regularization effect of second-order momentum in Adam is crucial to revitalize the weights that are dead due to the activation saturation in BNNs. We find that Adam, through its adaptive learning rate strategy, is better equipped to handle the rugged loss surface of BNNs and reaches a better optimum with higher generalization ability. Furthermore, we inspect the intriguing role of the real-valued weights in binary networks, and reveal the effect of weight decay on the stability and sluggishness of BNN optimization. Through extensive experiments and analysis, we derive a simple training scheme, building on existing Adam-based optimization, which achieves 70.5% top-1 accuracy on the ImageNet dataset using the same architecture as the state-of-the-art ReActNet (Liu et al., 2020) while achieving 1.1% higher accuracy. Code and models are available at https: //github.com/liuzechun/AdamBNN .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd9deae4-9298-48b5-845d-c9f1e153d2fdCited by top-tier papers15
- Training Transformers with 4-bit IntegersHaocheng Xi, Changhao Li, Jianfei Chen, Jun ZhuNeurIPS 2023 · 96 citations
- Oscillation-free Quantization for Low-bit Vision TransformersShih-Yang Liu, Zechun Liu, Kwang-Ting ChengICML 2023 · 63 citations
- BiViT: Extremely Compressed Binary Vision TransformersYefei He, Zhenyu Lou, Luoming Zhang, Jing Liu et al.ICCV 2023 · 44 citations
- LeHDC: learning-based hyperdimensional computing classifierShijin Duan, Yejia Liu, Shaolei Ren, Xiaolin XuDAC 2022 · 38 citations
- Binarized Neural Machine TranslationYichi Zhang, Ankush Garg, Yuan Cao, Lukasz Lew et al.NeurIPS 2023 · 21 citations
Builds on4
- Training binary neural networks with real-to-binary convolutionsBrais Martínez, Jing Yang, Adrian Bulat, Georgios TzimiropoulosICLR 2020 · 251 citations
- Binarizing MobileNet via Evolution-Based SearchingHai Phan, Zechun Liu, Dang Huynh, Marios Savvides et al.CVPR 2020
- S2-BNN: Bridging the Gap Between Self-Supervised Real and 1-Bit Neural Networks via Guided Distribution CalibrationZhiqiang Shen, Zechun Liu, Jie Qin, Lei Huang et al.CVPR 2021
- Forward and Backward Information Retention for Accurate Binary Neural NetworksHaotong Qin, Ruihao Gong, Xianglong Liu, Mingzhu Shen et al.CVPR 2020
Related papers
- Resilient Binary Neural NetworkSheng Xu, Yanjing Li, Teli Ma, Mingbao Lin et al.AAAI 2023 · 1 citation
- Understanding weight-magnitude hyperparameters in training binary networksJoris Quist, Yunqiang Li, Jan van GemertICLR 2023
- PokeBNN: A Binary Pursuit of Lightweight AccuracyYichi Zhang, Zhiru Zhang, Lukasz LewCVPR 2022 · 38 citations
- SA-BNN: State-Aware Binary Neural NetworkChunlei Liu, Peng Chen, Bohan Zhuang, Chunhua Shen et al.AAAI 2021 · 23 citations
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 267 citations
