COMET: A Novel Memory-Efficient Deep Learning Training Framework by Using Error-Bounded Lossy Compression
Sian Jin, Chengming Zhang, Xintong Jiang, Yunhe Feng, Hui Guan, Guanpeng Li, Shuaiwen Song, Dingwen Tao
Abstract
Deep neural networks (DNNs) are becoming increasingly deeper, wider, and non-linear due to the growing demands on prediction accuracy and analysis quality. Training wide and deep neural networks require large amounts of storage resources such as memory because the intermediate activation data must be saved in the memory during forward propagation and then restored for backward propagation. However, state-of-the-art accelerators such as GPUs are only equipped with very limited memory capacities due to hardware design constraints, which significantly limits the maximum batch size and hence performance speedup when training large-scale DNNs. Traditional memory saving techniques either suffer from performance overhead or are constrained by limited interconnect bandwidth or specific interconnect technology. In this paper, we propose a novel memory-efficient CNN training framework (called COMET) that leverages error-bounded lossy compression to significantly reduce the memory requirement for training in order to allow training larger models or to accelerate training. Different from the state-of-the-art solutions that adopt image-based lossy compressors (such as JPEG) to compress the activation data, our framework purposely adopts error-bounded lossy compression with a strict error-controlling mechanism. Specifically, we perform a theoretical analysis on the compression error propagation from the altered activation data to the gradients, and empirically investigate the impact of altered gradients over the training process. Based on these analyses, we optimize the errorbounded lossy compression and propose an adaptive error-bound control scheme for activation data compression. We evaluate our design against state-of-the-art solutions with five widely-adopted CNNs and ImageNet dataset. Experiments demonstrate that our proposed framework can significantly reduce the training memory consumption by up to 13.5× over the baseline training and 1.8× over another state-of-the-art compression-based framework, respectively, with little or no accuracy loss.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 313e6f06-ffb0-483a-9572-3bf5c389650aCited by top-tier papers6
- FZ-GPU: A Fast and High-Ratio Lossy Compressor for Scientific Computing Applications on GPUsBoyuan Zhang, Jiannan Tian, Sheng Di, Xiaodong Yu et al.HPDC 2023 · 27 citations
- DREW: Efficient Winograd CNN Inference with Deep ReuseRuofan Wu, Feng Zhang, Jiawei Guan, Zhen Zheng et al.WWW 2022 · 20 citations
- BladeDISC: Optimizing Dynamic Shape Machine Learning Workloads via Compiler ApproachZhen Zheng, Zaifeng Pan, Dalin Wang, Kai Zhu et al.SIGMOD 2024 · 14 citations
- CereSZ: Enabling and Scaling Error-bounded Lossy Compression on Cerebras CS-2Shihui Song, Yafan Huang, Peng Jiang, Xiaodong Yu et al.HPDC 2024 · 14 citations
- Concealing Compression-accelerated I/O for HPC Applications through In Situ Task SchedulingSian Jin, Sheng Di, Frédéric Vivien, Daoce Wang et al.EuroSys 2024 · 13 citations
Builds on4
- Cerebro: A Data System for Optimized Deep Learning Model SelectionSupun Nakandala, Yuhao Zhang, Arun KumarVLDB 2020 · 61 citations
- Buddy Compression: Enabling Larger Memory for Deep Learning and HPC Workloads on GPUsEsha Choukse, Michael B. Sullivan, Mike O'Connor, Mattan Erez et al.ISCA 2020 · 58 citations
- JPEG-ACT: Accelerating Deep Learning via Transform-based Lossy CompressionR. David Evans, Lufei Liu, Tor M. AamodtISCA 2020 · 47 citations
- Convolutional neural network training with distributed K-FACJ. Gregory Pauloski, Zhao Zhang, Lei Huang, Weijia Xu et al.SC 2020 · 26 citations
Related papers
- AC-GC: Lossy Activation Compression with Guaranteed ConvergenceR. David Evans, Tor M. AamodtNeurIPS 2021 · 38 citations
- INSTANT: Compressing Gradients and Activations for Resource-Efficient TrainingTuan-Kiet Doan, Trung-Hieu Tran, Enzo Tartaglione, Nikola Simidjievski et al.ICLR 2026
- GACT: Activation Compressed Training for Generic Network ArchitecturesXiaoxuan Liu, Lianmin Zheng, Dequan Wang, Yukuo Cen et al.ICML 2022 · 43 citations
- ALAM: Averaged Low-Precision Activation for Memory-Efficient Training of Transformer ModelsSunghyeon Woo, Sunwoo Lee, Dongsuk JeonICLR 2024 · 4 citations
- The Art of Losing to Win: Using Lossy Image Compression to Improve Data Loading in Deep Learning PipelinesLennart Behme, Saravanan Thirumuruganathan, Alireza Rezaei Mahdiraji, Jorge-Arnulfo Quiané-Ruiz et al.ICDE 2023 · 4 citations
