AC-GC: Lossy Activation Compression with Guaranteed Convergence
R. David Evans, Tor M. Aamodt
摘要
Parallel hardware devices (e.g., graphics processor units) have limited highbandwidth memory capacity. This negatively impacts the training of deep neural networks (DNNs) by increasing runtime and/or decreasing accuracy when reducing model and/or batch size to fit this capacity. Lossy compression is a promising approach to tackling memory capacity constraints, but prior approaches rely on hyperparameter search to achieve a suitable trade-off between convergence and compression, negating runtime benefits. In this paper we build upon recent developments on Stochastic Gradient Descent convergence to prove an upper bound on the expected loss increase when training with compressed activation storage. We then express activation compression error in terms of this bound, allowing the compression rate to adapt to training conditions automatically. The advantage of our approach, called AC-GC, over existing lossy compression frameworks is that, given a preset allowable increase in loss, significant compression without significant increase in error can be achieved with a single training run. When combined with error-bounded methods, AC-GC achieves 15.1× compression with an average accuracy change of 0.1% on text and image datasets. AC-GC functions on any model composed of the layers analyzed and, by avoiding compression rate search, reduces overall training time by 4.6× over SuccessiveHalving.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Scaling for Training Time and Post-hoc Out-of-distribution Detection EnhancementKai Xu, Rongyu Chen, Gianni Franchi, Angela YaoICLR 2024 · 被引用 81 次
- GACT: Activation Compressed Training for Generic Network ArchitecturesXiaoxuan Liu, Lianmin Zheng, Dequan Wang, Yukuo Cen 等ICML 2022 · 被引用 43 次
- Fine-tuning Language Models over Slow Networks using Activation Quantization with GuaranteesJue Wang, Binhang Yuan, Luka Rimanic, Yongjun He 等NeurIPS 2022 · 被引用 37 次
- Back Razor: Memory-Efficient Transfer Learning by Self-Sparsified BackpropagationZiyu Jiang, Xuxi Chen, Xueqin Huang, Xianzhi Du 等NeurIPS 2022 · 被引用 25 次
- TinyTrain: Resource-Aware Task-Adaptive Sparse Training of DNNs at the Data-Scarce EdgeYoung D. Kwon, Rui Li, Stylianos I. Venieris, Jagmohan Chauhan 等ICML 2024 · 被引用 25 次
它引用的顶会 Paper3
- Ultra-Low Precision 4-bit Training of Deep Neural NetworksXiao Sun, Naigang Wang, Chia-Yu Chen, Jiamin Ni 等NeurIPS 2020 · 被引用 227 次
- Sparse Weight Activation TrainingMd Aamir Raihan, Tor M. AamodtNeurIPS 2020 · 被引用 83 次
- JPEG-ACT: Accelerating Deep Learning via Transform-based Lossy CompressionR. David Evans, Lufei Liu, Tor M. AamodtISCA 2020 · 被引用 47 次
相关 Paper
- COMET: A Novel Memory-Efficient Deep Learning Training Framework by Using Error-Bounded Lossy CompressionSian Jin, Chengming Zhang, Xintong Jiang, Yunhe Feng 等VLDB 2022 · 被引用 39 次
- The Art of Losing to Win: Using Lossy Image Compression to Improve Data Loading in Deep Learning PipelinesLennart Behme, Saravanan Thirumuruganathan, Alireza Rezaei Mahdiraji, Jorge-Arnulfo Quiané-Ruiz 等ICDE 2023 · 被引用 4 次
- INSTANT: Compressing Gradients and Activations for Resource-Efficient TrainingTuan-Kiet Doan, Trung-Hieu Tran, Enzo Tartaglione, Nikola Simidjievski 等ICLR 2026
- DIVISION: Memory Efficient Training via Dual Activation PrecisionGuanchu Wang, Zirui Liu, Zhimeng Jiang, Ninghao Liu 等ICML 2023 · 被引用 4 次
- On the Discrepancy between the Theoretical Analysis and Practical Implementations of Compressed Communication for Distributed Deep LearningAritra Dutta, El Houcine Bergou, Ahmed M. Abdelmoniem, Chen-Yu Ho 等AAAI 2020
