Neural gradients are near-lognormal: improved quantized and sparse training
Brian Chmiel, Liad Ben-Uri, Moran Shkolnik, Elad Hoffer, Ron Banner, Daniel Soudry
摘要
While training can mostly be accelerated by reducing the time needed to propagate neural gradients back throughout the model, most previous works focus on the quantization/pruning of weights and activations. These methods are often not applicable to neural gradients, which have very different statistical properties. Distinguished from weights and activations, we find that the distribution of neural gradients is approximately lognormal. Considering this, we suggest two closed-form analytical methods to reduce the computational and memory burdens of neural gradients. The first method optimizes the floating-point format and scale of the gradients. The second method accurately sets sparsity thresholds for gradient pruning. Each method achieves state-of-the-art results on ImageNet. To the best of our knowledge, this paper is the first to (1) quantize the gradients to 6-bit floating-point formats, or (2) achieve up to 85% gradient sparsity -- in each case without accuracy degradation. Reference implementation accompanies the paper.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- DRIVE: One-bit Distributed Mean EstimationShay Vargaftik, Ran Ben-Basat, Amit Portnoy, Gal Mendelson 等NeurIPS 2021 · 被引用 82 次
- FAST: DNN Training Under Variable Precision Block Floating Point with Stochastic RoundingSai Qian Zhang, Bradley McDanel, H. T. KungHPCA 2022 · 被引用 68 次
- EDEN: Communication-Efficient and Robust Distributed Mean Estimation for Federated LearningShay Vargaftik, Ran Ben Basat, Amit Portnoy, Gal Mendelson 等ICML 2022 · 被引用 64 次
- The future is log-Gaussian: ResNets and their infinite-depth-and-width limit at initializationMufan Bill Li, Mihai Nica, Daniel M. RoyNeurIPS 2021 · 被引用 41 次
- Differentiable Weightless Neural NetworksAlan Tendler Leibel Bacellar, Zachary Susskind, Maurício Breternitz Jr., Eugene John 等ICML 2024 · 被引用 34 次
它引用的顶会 Paper2
相关 Paper
- Accurate Neural Training with 4-bit Matrix Multiplications at Standard FormatsBrian Chmiel, Ron Banner, Elad Hoffer, Hilla Ben-Yaacov 等ICLR 2023 · 被引用 6 次
- Minimum Variance Unbiased N: M Sparsity for the Neural GradientsBrian Chmiel, Itay Hubara, Ron Banner, Daniel SoudryICLR 2023
- Distribution Adaptive INT8 Quantization for Training CNNsKang Zhao, Sida Huang, Pan Pan, Yinghan Li 等AAAI 2021 · 被引用 86 次
- Few-bit Backward: Quantized Gradients of Activation Functions for Memory Footprint ReductionGeorgii Sergeevich Novikov, Daniel Bershatsky, Julia Gusak, Alex Shonenkov 等ICML 2023 · 被引用 19 次
- SparseTrain: Exploiting Dataflow Sparsity for Efficient Convolutional Neural Networks TrainingPengcheng Dai, Jianlei Yang, Xucheng Ye, Xingzhou Cheng 等DAC 2020 · 被引用 27 次
