How Do Adam and Training Strategies Help BNNs Optimization
Zechun Liu, Zhiqiang Shen, Shichao Li, Koen Helwegen, Dong Huang, Kwang-Ting Cheng
摘要
The best performing Binary Neural Networks (BNNs) are usually attained using Adam optimization and its multi-step training variants (Rastegari et al., 2016; Liu et al., 2020) . However, to the best of our knowledge, few studies explore the fundamental reasons why Adam is superior to other optimizers like SGD for BNN optimization or provide analytical explanations that support specific training strategies. To address this, in this paper we first investigate the trajectories of gradients and weights in BNNs during the training process. We show the regularization effect of second-order momentum in Adam is crucial to revitalize the weights that are dead due to the activation saturation in BNNs. We find that Adam, through its adaptive learning rate strategy, is better equipped to handle the rugged loss surface of BNNs and reaches a better optimum with higher generalization ability. Furthermore, we inspect the intriguing role of the real-valued weights in binary networks, and reveal the effect of weight decay on the stability and sluggishness of BNN optimization. Through extensive experiments and analysis, we derive a simple training scheme, building on existing Adam-based optimization, which achieves 70.5% top-1 accuracy on the ImageNet dataset using the same architecture as the state-of-the-art ReActNet (Liu et al., 2020) while achieving 1.1% higher accuracy. Code and models are available at https: //github.com/liuzechun/AdamBNN .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Training Transformers with 4-bit IntegersHaocheng Xi, Changhao Li, Jianfei Chen, Jun ZhuNeurIPS 2023 · 被引用 96 次
- Oscillation-free Quantization for Low-bit Vision TransformersShih-Yang Liu, Zechun Liu, Kwang-Ting ChengICML 2023 · 被引用 63 次
- BiViT: Extremely Compressed Binary Vision TransformersYefei He, Zhenyu Lou, Luoming Zhang, Jing Liu 等ICCV 2023 · 被引用 44 次
- LeHDC: learning-based hyperdimensional computing classifierShijin Duan, Yejia Liu, Shaolei Ren, Xiaolin XuDAC 2022 · 被引用 38 次
- Binarized Neural Machine TranslationYichi Zhang, Ankush Garg, Yuan Cao, Lukasz Lew 等NeurIPS 2023 · 被引用 21 次
它引用的顶会 Paper4
- Training binary neural networks with real-to-binary convolutionsBrais Martínez, Jing Yang, Adrian Bulat, Georgios TzimiropoulosICLR 2020 · 被引用 251 次
- Binarizing MobileNet via Evolution-Based SearchingHai Phan, Zechun Liu, Dang Huynh, Marios Savvides 等CVPR 2020
- S2-BNN: Bridging the Gap Between Self-Supervised Real and 1-Bit Neural Networks via Guided Distribution CalibrationZhiqiang Shen, Zechun Liu, Jie Qin, Lei Huang 等CVPR 2021
- Forward and Backward Information Retention for Accurate Binary Neural NetworksHaotong Qin, Ruihao Gong, Xianglong Liu, Mingzhu Shen 等CVPR 2020
相关 Paper
- Resilient Binary Neural NetworkSheng Xu, Yanjing Li, Teli Ma, Mingbao Lin 等AAAI 2023 · 被引用 1 次
- Understanding weight-magnitude hyperparameters in training binary networksJoris Quist, Yunqiang Li, Jan van GemertICLR 2023
- PokeBNN: A Binary Pursuit of Lightweight AccuracyYichi Zhang, Zhiru Zhang, Lukasz LewCVPR 2022 · 被引用 38 次
- SA-BNN: State-Aware Binary Neural NetworkChunlei Liu, Peng Chen, Bohan Zhuang, Chunhua Shen 等AAAI 2021 · 被引用 23 次
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 被引用 267 次
