Few-bit Backward: Quantized Gradients of Activation Functions for Memory Footprint Reduction
Georgii Sergeevich Novikov, Daniel Bershatsky, Julia Gusak, Alex Shonenkov, Denis Valerievich Dimitrov, Ivan V. Oseledets
摘要
Memory footprint is one of the main limiting factors for large neural network training. In backpropagation, one needs to store the input to each operation in the computational graph. Every modern neural network model has quite a few pointwise nonlinearities in its architecture, and such operation induces additional memory costs which -as we show -can be significantly reduced by quantization of the gradients. We propose a systematic approach to compute optimal quantization of the retained gradients of the pointwise nonlinear functions with only a few bits per each element. We show that such approximation can be achieved by computing optimal piecewiseconstant approximation of the derivative of the activation function, which can be done by dynamic programming. The drop-in replacements are implemented for all popular nonlinearities and can be used in any existing pipeline. We confirm the memory reduction and the same convergence on several open benchmarks. The proposed approximation divides all values into several bins and saves only its corresponding bin indices instead of storing all values. This is a lossy compresion, but the additional noise introduced by it is negligible as we will show on several benchmarks 4. Main contributions of our paper are: • We propose new approximate backward computation
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- 8-bit Optimizers via Block-wise QuantizationTim Dettmers, Mike Lewis, Sam Shleifer, Luke ZettlemoyerICLR 2022 · 被引用 457 次
- Understanding and Overcoming the Challenges of Efficient Transformer QuantizationYelysei Bondarenko, Markus Nagel, Tijmen BlankevoortEMNLP 2021 · 被引用 74 次
- Efficient Combination of Rematerialization and Offloading for Training DNNsOlivier Beaumont, Lionel Eyraud-Dubois, Alena ShilovaNeurIPS 2021 · 被引用 69 次
相关 Paper
- Accurate Neural Training with 4-bit Matrix Multiplications at Standard FormatsBrian Chmiel, Ron Banner, Elad Hoffer, Hilla Ben-Yaacov 等ICLR 2023 · 被引用 6 次
- Optimal and Approximate Adaptive Stochastic QuantizationRan Ben-Basat, Yaniv Ben-Itzhak, Michael Mitzenmacher, Shay VargaftikNeurIPS 2024 · 被引用 12 次
- DIVISION: Memory Efficient Training via Dual Activation PrecisionGuanchu Wang, Zirui Liu, Zhimeng Jiang, Ninghao Liu 等ICML 2023 · 被引用 4 次
- Towards Cheaper Inference in Deep Networks with Lower Bit-Width AccumulatorsYaniv Blumenfeld, Itay Hubara, Daniel SoudryICLR 2024 · 被引用 5 次
- Indirect Stochastic Gradient Quantization and Its Application in Distributed Deep LearningAfshin Abdi, Faramarz FekriAAAI 2020 · 被引用 5 次
