Few-bit Backward: Quantized Gradients of Activation Functions for Memory Footprint Reduction
Georgii Sergeevich Novikov, Daniel Bershatsky, Julia Gusak, Alex Shonenkov, Denis Valerievich Dimitrov, Ivan V. Oseledets
Abstract
Memory footprint is one of the main limiting factors for large neural network training. In backpropagation, one needs to store the input to each operation in the computational graph. Every modern neural network model has quite a few pointwise nonlinearities in its architecture, and such operation induces additional memory costs which -as we show -can be significantly reduced by quantization of the gradients. We propose a systematic approach to compute optimal quantization of the retained gradients of the pointwise nonlinear functions with only a few bits per each element. We show that such approximation can be achieved by computing optimal piecewiseconstant approximation of the derivative of the activation function, which can be done by dynamic programming. The drop-in replacements are implemented for all popular nonlinearities and can be used in any existing pipeline. We confirm the memory reduction and the same convergence on several open benchmarks. The proposed approximation divides all values into several bins and saves only its corresponding bin indices instead of storing all values. This is a lossy compresion, but the additional noise introduced by it is negligible as we will show on several benchmarks 4. Main contributions of our paper are: • We propose new approximate backward computation
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- 8-bit Optimizers via Block-wise QuantizationTim Dettmers, Mike Lewis, Sam Shleifer, Luke ZettlemoyerICLR 2022 · 457 citations
- Understanding and Overcoming the Challenges of Efficient Transformer QuantizationYelysei Bondarenko, Markus Nagel, Tijmen BlankevoortEMNLP 2021 · 74 citations
- Efficient Combination of Rematerialization and Offloading for Training DNNsOlivier Beaumont, Lionel Eyraud-Dubois, Alena ShilovaNeurIPS 2021 · 69 citations
Related papers
- Accurate Neural Training with 4-bit Matrix Multiplications at Standard FormatsBrian Chmiel, Ron Banner, Elad Hoffer, Hilla Ben-Yaacov et al.ICLR 2023 · 6 citations
- Optimal and Approximate Adaptive Stochastic QuantizationRan Ben-Basat, Yaniv Ben-Itzhak, Michael Mitzenmacher, Shay VargaftikNeurIPS 2024 · 12 citations
- DIVISION: Memory Efficient Training via Dual Activation PrecisionGuanchu Wang, Zirui Liu, Zhimeng Jiang, Ninghao Liu et al.ICML 2023 · 4 citations
- Towards Cheaper Inference in Deep Networks with Lower Bit-Width AccumulatorsYaniv Blumenfeld, Itay Hubara, Daniel SoudryICLR 2024 · 5 citations
- Indirect Stochastic Gradient Quantization and Its Application in Distributed Deep LearningAfshin Abdi, Faramarz FekriAAAI 2020 · 5 citations
