Training binary neural networks with real-to-binary convolutions
Brais Martínez, Jing Yang, Adrian Bulat, Georgios Tzimiropoulos
Abstract
This paper shows how to train binary networks to within a few percent points ( 3-5 %) of the full precision counterpart with a negligible increase in the computational cost. In particular, we first show how to build a strong baseline, which already achieves state-of-the-art accuracy, by combining recently proposed advances, and carefully tuning the optimization procedure. Secondly, we show that by attempting to minimize the discrepancy between the output of the binary and the corresponding real-valued convolution additional significant accuracy gains can be obtained. We materialize this idea in two complementary ways: (1) with a loss function, during training, by matching the spatial attention maps computed at the output of the binary and real-valued convolutions, and (2) in data-driven manner, by using the real-valued activations being available during inference prior to the binarization process for re-scaling the activations right after the binary convolution. Finally, we show that, when putting all of our improvements together, the resulting model reduces the gap to its real-valued counterpart to less than 3% and 5% top-1 error on CIFAR-100 and ImageNet, respectively, when using a ResNet-18 architecture.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 500e6ae1-3782-4220-b8be-7670791e7b5fCited by top-tier papers27
- BiBERT: Accurate Fully Binarized BERTHaotong Qin, Yifu Ding, Mingyuan Zhang, Qinghua Yan et al.ICLR 2022 · 121 citations
- How Do Adam and Training Strategies Help BNNs OptimizationZechun Liu, Zhiqiang Shen, Shichao Li, Koen Helwegen et al.ICML 2021 · 100 citations
- BiT: Robustly Binarized Multi-distilled TransformerZechun Liu, Barlas Oguz, Aasish Pappu, Lin Xiao et al.NeurIPS 2022 · 93 citations
- Is Label Smoothing Truly Incompatible with Knowledge Distillation: An Empirical StudyZhiqiang Shen, Zechun Liu, Dejia Xu, Zitian Chen et al.ICLR 2021 · 83 citations
- BiPointNet: Binary Neural Network for Point CloudsHaotong Qin, Zhongang Cai, Mingyuan Zhang, Yifu Ding et al.ICLR 2021 · 54 citations
Builds on1
Related papers
- High-Capacity Expert Binary NetworksAdrian Bulat, Brais Martínez, Georgios TzimiropoulosICLR 2021 · 29 citations
- Sparsity-Inducing Binarized Neural NetworksPeisong Wang, Xiangyu He, Gang Li, Tianli Zhao et al.AAAI 2020 · 60 citations
- Fast and Accurate Binary Neural Networks Based on Depth-Width ReshapingPing Xue, Yang Lu, Jingfei Chang, Xing Wei et al.AAAI 2023 · 3 citations
- Training Binary Neural Network without Batch Normalization for Image Super-ResolutionXinrui Jiang, Nannan Wang, Jingwei Xin, Keyu Li et al.AAAI 2021 · 52 citations
- SA-BNN: State-Aware Binary Neural NetworkChunlei Liu, Peng Chen, Bohan Zhuang, Chunhua Shen et al.AAAI 2021 · 23 citations
