Training Binary Neural Networks using the Bayesian Learning Rule
Xiangming Meng, Roman Bachmann, Mohammad Emtiyaz Khan
摘要
Neural networks with binary weights are computation-efficient and hardware-friendly, but their training is challenging because it involves a discrete optimization problem. Surprisingly, ignoring the discrete nature of the problem and using gradient-based methods, such as the Straight-Through Estimator, still works well in practice. This raises the question: are there principled approaches which justify such methods? In this paper, we propose such an approach using the Bayesian learning rule. The rule, when applied to estimate a Bernoulli distribution over the binary weights, results in an algorithm which justifies some of the algorithmic choices made by the previous approaches. The algorithm not only obtains state-of-the-art performance, but also enables uncertainty estimation for continual learning to avoid catastrophic forgetting. Our work provides a principled approach for training binary neural networks which justifies and extends existing approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- AutoReP: Automatic ReLU Replacement for Fast Private Network InferenceHongwu Peng, Shaoyi Huang, Tong Zhou, Yukui Luo 等ICCV 2023 · 被引用 44 次
- Low-Precision Stochastic Gradient Langevin DynamicsRuqi Zhang, Andrew Gordon Wilson, Christopher De SaICML 2022 · 被引用 18 次
- Unsupervised Representation Learning for Binary Networks by Joint Classifier LearningDahyun Kim, Jonghyun ChoiCVPR 2022 · 被引用 6 次
- Training Binary Neural Networks via Gaussian Variational Inference and Low-Rank Semidefinite ProgrammingLorenzo Orecchia, Jiawei Hu, Xue He, Wang Mark 等NeurIPS 2024 · 被引用 4 次
- Understanding weight-magnitude hyperparameters in training binary networksJoris Quist, Yunqiang Li, Jan van GemertICLR 2023
它引用的顶会 Paper1
相关 Paper
- Path Sample-Analytic Gradient Estimators for Stochastic Binary NetworksAlexander Shekhovtsov, Viktor Yanush, Boris FlachNeurIPS 2020 · 被引用 14 次
- AdaSTE: An Adaptive Straight-Through Estimator to Train Binary Neural NetworksHuu Le, Rasmus Kjær Høier, Che-Tsung Lin, Christopher ZachCVPR 2022
- Uncertainty-guided Continual Learning with Bayesian Neural NetworksSayna Ebrahimi, Mohamed Elhoseiny, Trevor Darrell, Marcus RohrbachICLR 2020 · 被引用 211 次
- Injecting Logical Constraints into Neural Networks via Straight-Through EstimatorsZhun Yang, Joohyung Lee, Chiyoun ParkICML 2022 · 被引用 26 次
- Joint Inference for Neural Network Depth and Dropout RegularizationKishan K. C., Rui Li, Mahdi GilanyNeurIPS 2021 · 被引用 13 次
