Low-Precision Stochastic Gradient Langevin Dynamics
Ruqi Zhang, Andrew Gordon Wilson, Christopher De Sa
摘要
While low-precision optimization has been widely used to accelerate deep learning, low-precision sampling remains largely unexplored. As a consequence, sampling is simply infeasible in many large-scale scenarios, despite providing remarkable benefits to generalization and uncertainty estimation for neural networks. In this paper, we provide the first study of low-precision Stochastic Gradient Langevin Dynamics (SGLD), showing that its costs can be significantly reduced without sacrificing performance, due to its intrinsic ability to handle system noise. We prove that the convergence of low-precision SGLD with full-precision gradient accumulators is less affected by the quantization error than its SGD counterpart in the strongly convex setting. To further enable low-precision gradient accumulators, we develop a new quantization function for SGLD that preserves the variance in each update step. We demonstrate that low-precision SGLD achieves comparable performance to full-precision SGLD with only 8 bits on a variety of deep learning tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Quantification of Uncertainty with Adversarial ModelsKajetan Schweighofer, Lukas Aichberger, Mykyta Ielanskyi, Günter Klambauer 等NeurIPS 2023 · 被引用 37 次
- BOLD: Boolean Logic Deep LearningVan Minh Nguyen, Cristian Ocampo-Blandon, Aymen Askri, Louis Leconte 等NeurIPS 2024 · 被引用 4 次
- Log-Normal Multiplicative Dynamics for Stable Low-Precision Deep LearningKeigo Nishida, Eren Mehmet KIRAL, Kenichi Bannai, Mohammad Emtiyaz Khan 等ICML 2026 · 被引用 2 次
- Ex Uno Pluria: Insights on Ensembling in Low Precision Number SystemsGiung Nam, Juho LeeNeurIPS 2024 · 被引用 2 次
它引用的顶会 Paper6
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy 等ICLR 2020 · 被引用 1,037 次
- Cyclical Stochastic Gradient MCMC for Bayesian Deep LearningRuqi Zhang, Chunyuan Li, Jianyi Zhang, Changyou Chen 等ICLR 2020 · 被引用 292 次
- Ultra-Low Precision 4-bit Training of Deep Neural NetworksXiao Sun, Naigang Wang, Chia-Yu Chen, Jiamin Ni 等NeurIPS 2020 · 被引用 227 次
- Bayesian Bits: Unifying Quantization and PruningMart van Baalen, Christos Louizos, Markus Nagel, Rana Ali Amjad 等NeurIPS 2020 · 被引用 149 次
- Training Binary Neural Networks using the Bayesian Learning RuleXiangming Meng, Roman Bachmann, Mohammad Emtiyaz KhanICML 2020 · 被引用 47 次
相关 Paper
- Towards Cheaper Inference in Deep Networks with Lower Bit-Width AccumulatorsYaniv Blumenfeld, Itay Hubara, Daniel SoudryICLR 2024 · 被引用 5 次
- The Marginal Value of Momentum for Small Learning Rate SGDRunzhe Wang, Sadhika Malladi, Tianhao Wang, Kaifeng Lyu 等ICLR 2024 · 被引用 14 次
- Flatness-Aware Stochastic Gradient Langevin DynamicsStefano Bruno, Youngsik Hwang, JaeHyeon An, Sotirios Sabanis 等ICML 2026
- Fractional Underdamped Langevin Dynamics: Retargeting SGD with Momentum under Heavy-Tailed Gradient NoiseUmut Simsekli, Lingjiong Zhu, Yee Whye Teh, Mert GürbüzbalabanICML 2020 · 被引用 58 次
- A Contour Stochastic Gradient Langevin Dynamics Algorithm for Simulations of Multi-modal DistributionsWei Deng, Guang Lin, Faming LiangNeurIPS 2020 · 被引用 37 次
