On Efficient Constructions of Checkpoints
Yu Chen, Zhenming Liu, Bin Ren, Xin Jin
Abstract
Efficient construction of checkpoints/snapshots is a critical tool for training and diagnosing deep learning models. In this paper, we propose a lossy compression scheme for checkpoint constructions (called LC-Checkpoint). LC-Checkpoint simultaneously maximizes the compression rate and optimizes the recovery speed, under the assumption that SGD is used to train the model. LC-Checkpointuses quantization and priority promotion to store the most crucial information for SGD to recover, and then uses a Huffman coding to leverage the non-uniform distribution of the gradient scales. Our extensive experiments show that LC-Checkpoint achieves a compression rate up to and recovery speedup up to over a state-of-the-art algorithm (SCAR).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- PipeSwitch: Fast Pipelined Context Switching for Deep Learning ApplicationsZhihao Bai, Zhen Zhang, Yibo Zhu, Xin JinOSDI 2020 · 152 citations
- GEMINI: Fast Failure Recovery in Distributed Training with In-Memory CheckpointsZhuang Wang, Zhen Jia, Shuai Zheng, Zhen Zhang et al.SOSP 2023 · 61 citations
- ExCP: Extreme LLM Checkpoint Compression via Weight-Momentum Joint ShrinkingWenshuo Li, Xinghao Chen, Han Shu, Yehui Tang et al.ICML 2024 · 11 citations
- Efficient Fault Tolerance for Recommendation Model Training via Erasure CodingTianyu Zhang, Kaige Liu, Jack Kosaian, Juncheng Yang et al.VLDB 2023 · 10 citations
- FM-Delta: Lossless Compression for Storing Massive Fine-tuned Foundation ModelsWanyi Ning, Jingyu Wang, Qi Qi, Mengde Zhu et al.NeurIPS 2024 · 10 citations
Related papers
- Indirect Stochastic Gradient Quantization and Its Application in Distributed Deep LearningAfshin Abdi, Faramarz FekriAAAI 2020 · 5 citations
- Towards Lossless Memory-efficient Training of Spiking Neural Networks via Gradient Checkpointing and Spike CompressionYifan Huang, Wei Fang, Zecheng Hao, Zhengyu Ma et al.ICLR 2026
- SK-Gradient: Efficient Communication for Distributed Machine Learning with Data SketchJie Gui, Yuchen Song, Zezhou Wang, Chenhong He et al.ICDE 2023 · 9 citations
- Lossy and Lossless (L2) Post-training Model Size CompressionYumeng Shi, Shihao Bai, Xiuying Wei, Ruihao Gong et al.ICCV 2023 · 5 citations
- Communication Efficient SGD via Gradient Sampling With Bayes PriorLiuyihan Song, Kang Zhao, Pan Pan, Yu Liu et al.CVPR 2021
