Low-Precision Stochastic Gradient Langevin Dynamics
Ruqi Zhang, Andrew Gordon Wilson, Christopher De Sa
Abstract
While low-precision optimization has been widely used to accelerate deep learning, low-precision sampling remains largely unexplored. As a consequence, sampling is simply infeasible in many large-scale scenarios, despite providing remarkable benefits to generalization and uncertainty estimation for neural networks. In this paper, we provide the first study of low-precision Stochastic Gradient Langevin Dynamics (SGLD), showing that its costs can be significantly reduced without sacrificing performance, due to its intrinsic ability to handle system noise. We prove that the convergence of low-precision SGLD with full-precision gradient accumulators is less affected by the quantization error than its SGD counterpart in the strongly convex setting. To further enable low-precision gradient accumulators, we develop a new quantization function for SGLD that preserves the variance in each update step. We demonstrate that low-precision SGLD achieves comparable performance to full-precision SGLD with only 8 bits on a variety of deep learning tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0d0a5e90-8cc8-4adf-910d-290efa83e639Cited by top-tier papers4
- Quantification of Uncertainty with Adversarial ModelsKajetan Schweighofer, Lukas Aichberger, Mykyta Ielanskyi, Günter Klambauer et al.NeurIPS 2023 · 37 citations
- BOLD: Boolean Logic Deep LearningVan Minh Nguyen, Cristian Ocampo-Blandon, Aymen Askri, Louis Leconte et al.NeurIPS 2024 · 4 citations
- Log-Normal Multiplicative Dynamics for Stable Low-Precision Deep LearningKeigo Nishida, Eren Mehmet KIRAL, Kenichi Bannai, Mohammad Emtiyaz Khan et al.ICML 2026 · 2 citations
- Ex Uno Pluria: Insights on Ensembling in Low Precision Number SystemsGiung Nam, Juho LeeNeurIPS 2024 · 2 citations
Builds on6
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy et al.ICLR 2020 · 1,037 citations
- Cyclical Stochastic Gradient MCMC for Bayesian Deep LearningRuqi Zhang, Chunyuan Li, Jianyi Zhang, Changyou Chen et al.ICLR 2020 · 292 citations
- Ultra-Low Precision 4-bit Training of Deep Neural NetworksXiao Sun, Naigang Wang, Chia-Yu Chen, Jiamin Ni et al.NeurIPS 2020 · 227 citations
- Bayesian Bits: Unifying Quantization and PruningMart van Baalen, Christos Louizos, Markus Nagel, Rana Ali Amjad et al.NeurIPS 2020 · 149 citations
- Training Binary Neural Networks using the Bayesian Learning RuleXiangming Meng, Roman Bachmann, Mohammad Emtiyaz KhanICML 2020 · 47 citations
Related papers
- Towards Cheaper Inference in Deep Networks with Lower Bit-Width AccumulatorsYaniv Blumenfeld, Itay Hubara, Daniel SoudryICLR 2024 · 5 citations
- The Marginal Value of Momentum for Small Learning Rate SGDRunzhe Wang, Sadhika Malladi, Tianhao Wang, Kaifeng Lyu et al.ICLR 2024 · 14 citations
- Flatness-Aware Stochastic Gradient Langevin DynamicsStefano Bruno, Youngsik Hwang, JaeHyeon An, Sotirios Sabanis et al.ICML 2026
- Fractional Underdamped Langevin Dynamics: Retargeting SGD with Momentum under Heavy-Tailed Gradient NoiseUmut Simsekli, Lingjiong Zhu, Yee Whye Teh, Mert GürbüzbalabanICML 2020 · 58 citations
- A Contour Stochastic Gradient Langevin Dynamics Algorithm for Simulations of Multi-modal DistributionsWei Deng, Guang Lin, Faming LiangNeurIPS 2020 · 37 citations
