Estimator Meets Equilibrium Perspective: A Rectified Straight Through Estimator for Binary Neural Networks Training
Xiao-Ming Wu, Dian Zheng, Zuhao Liu, Wei-Shi Zheng
摘要
Binarization of neural networks is a dominant paradigm in neural networks compression. The pioneering work BinaryConnect uses Straight Through Estimator (STE) to mimic the gradients of the sign function, but it also causes the crucial inconsistency problem. Most of the previous methods design different estimators instead of STE to mitigate it. However, they ignore the fact that when reducing the estimating error, the gradient stability will decrease concomitantly. These highly divergent gradients will harm the model training and increase the risk of gradient vanishing and gradient exploding. To fully take the gradient stability into consideration, we present a new perspective to the BNNs training, regarding it as the equilibrium between the estimating error and the gradient stability. In this view, we firstly design two indicators to quantitatively demonstrate the equilibrium phenomenon. In addition, in order to balance the estimating error and the gradient stability well, we revise the original straight through estimator and propose a power function based estimator, Rectified Straight Through Estimator (ReSTE for short). Comparing to other estimators, ReSTE is rational and capable of flexibly balancing the estimating error with the gradient stability. Extensive experiments on CIFAR-10 and ImageNet datasets show that ReSTE has excellent performance and surpasses the state-of-the-art methods without any auxiliary modules or losses.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- BiDM: Pushing the Limit of Quantization for Diffusion ModelsXingyu Zheng, Xianglong Liu, Yichen Bian, Xudong Ma 等NeurIPS 2024 · 被引用 12 次
- BiPer: Binary Neural Networks Using a Periodic FunctionEdwin Vargas, Claudia V. Correa P., Carlos Hinojosa, Henry ArguelloCVPR 2024 · 被引用 10 次
- Binarized Low-Light Raw Video EnhancementGengchen Zhang, Yulun Zhang, Xin Yuan, Ying FuCVPR 2024 · 被引用 10 次
- Training Binary Neural Networks via Gaussian Variational Inference and Low-Rank Semidefinite ProgrammingLorenzo Orecchia, Jiawei Hu, Xue He, Wang Mark 等NeurIPS 2024 · 被引用 4 次
- Viperson: Flexibly Generating Virtual Identity for Person Re-IdentificationXiao-Wen Zhang, Delong Zhang, Yi-Xing Peng, Zhi Ouyang 等ICCV 2025 · 被引用 2 次
它引用的顶会 Paper7
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu 等CVPR 2022 · 被引用 835 次
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural NetworksRuihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li 等ICCV 2019 · 被引用 540 次
- ResRep: Lossless CNN Pruning via Decoupling Remembering and ForgettingXiaohan Ding, Tianxiang Hao, Jianchao Tan, Ji Liu 等ICCV 2021 · 被引用 202 次
- Rotated Binary Neural NetworkMingbao Lin, Rongrong Ji, Zihan Xu, Baochang Zhang 等NeurIPS 2020 · 被引用 161 次
- Learning Frequency Domain Approximation for Binary Neural NetworksYixing Xu, Kai Han, Chang Xu, Yehui Tang 等NeurIPS 2021 · 被引用 64 次
相关 Paper
- SURGE: Surrogate Gradient Adaptation in Binary Neural NetworksHaoyu Huang, Boyu Liu, Linlin Yang, Yanjing Li 等ICML 2026
- Network Quantization With Element-Wise Gradient ScalingJunghyup Lee, Dohyung Kim, Bumsub HamCVPR 2021
- Compacting Binary Neural Networks by Sparse Kernel SelectionYikai Wang, Wenbing Huang, Yinpeng Dong, Fuchun Sun 等CVPR 2023
- Understanding Neural Network Binarization with Forward and Backward Proximal QuantizersYiwei Lu, Yaoliang Yu, Xinlin Li, Vahid Partovi NiaNeurIPS 2023 · 被引用 5 次
- ReCU: Reviving the Dead Weights in Binary Neural NetworksZihan Xu, Mingbao Lin, Jianzhuang Liu, Jie Chen 等ICCV 2021 · 被引用 102 次
