Understanding Neural Network Binarization with Forward and Backward Proximal Quantizers
Yiwei Lu, Yaoliang Yu, Xinlin Li, Vahid Partovi Nia
Abstract
In neural network binarization, BinaryConnect (BC) and its variants are considered the standard. These methods apply the sign function in their forward pass and their respective gradients are backpropagated to update the weights. However, the derivative of the sign function is zero whenever defined, which consequently freezes training. Therefore, implementations of BC (e.g., BNN) usually replace the derivative of sign in the backward computation with identity or other approximate gradient alternatives. Although such practice works well empirically, it is largely a heuristic or ''training trick.'' We aim at shedding some light on these training tricks from the optimization perspective. Building from existing theory on ProxConnect (PC, a generalization of BC), we (1) equip PC with different forward-backward quantizers and obtain ProxConnect++ (PC++) that includes existing binarization techniques as special cases; (2) derive a principled way to synthesize forward-backward quantizers with automatic theoretical guarantees; (3) illustrate our theory by proposing an enhanced binarization algorithm BNN++; (4) conduct image classification experiments on CNNs and vision transformers, and empirically verify that BNN++ generally achieves competitive results on binarizing these models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a9481fac-c2e0-442e-93bd-2badaec32440Cited by top-tier papers3
- S2NN: Sub-bit Spiking Neural NetworksWenjie Wei, Malu Zhang, Jieyuan Zhang, Ammar Belatreche et al.NeurIPS 2025 · 1 citation
- BHViT: Binarized Hybrid Vision TransformerTian Gao, Yu Zhang, Zhiyuan Zhang, Huajun Liu et al.CVPR 2025
- PARQ: Piecewise-Affine Regularized QuantizationLisa Jin, Jianhao Ma, Zechun Liu, Andrey Gromov et al.ICML 2025
Builds on28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
Related papers
- Demystifying and Generalizing BinaryConnectTim Dockhorn, Yaoliang Yu, Eyyüb Sari, Mahdi Zolnouri et al.NeurIPS 2021 · 14 citations
- Sparsity-Inducing Binarized Neural NetworksPeisong Wang, Xiangyu He, Gang Li, Tianli Zhao et al.AAAI 2020 · 60 citations
- Estimator Meets Equilibrium Perspective: A Rectified Straight Through Estimator for Binary Neural Networks TrainingXiao-Ming Wu, Dian Zheng, Zuhao Liu, Wei-Shi ZhengICCV 2023 · 28 citations
- BiPer: Binary Neural Networks Using a Periodic FunctionEdwin Vargas, Claudia V. Correa P., Carlos Hinojosa, Henry ArguelloCVPR 2024 · 10 citations
- Rotated Binary Neural NetworkMingbao Lin, Rongrong Ji, Zihan Xu, Baochang Zhang et al.NeurIPS 2020 · 161 citations
