On Efficient Constructions of Checkpoints
Yu Chen, Zhenming Liu, Bin Ren, Xin Jin
摘要
Efficient construction of checkpoints/snapshots is a critical tool for training and diagnosing deep learning models. In this paper, we propose a lossy compression scheme for checkpoint constructions (called LC-Checkpoint). LC-Checkpoint simultaneously maximizes the compression rate and optimizes the recovery speed, under the assumption that SGD is used to train the model. LC-Checkpointuses quantization and priority promotion to store the most crucial information for SGD to recover, and then uses a Huffman coding to leverage the non-uniform distribution of the gradient scales. Our extensive experiments show that LC-Checkpoint achieves a compression rate up to and recovery speedup up to over a state-of-the-art algorithm (SCAR).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- PipeSwitch: Fast Pipelined Context Switching for Deep Learning ApplicationsZhihao Bai, Zhen Zhang, Yibo Zhu, Xin JinOSDI 2020 · 被引用 152 次
- GEMINI: Fast Failure Recovery in Distributed Training with In-Memory CheckpointsZhuang Wang, Zhen Jia, Shuai Zheng, Zhen Zhang 等SOSP 2023 · 被引用 61 次
- ExCP: Extreme LLM Checkpoint Compression via Weight-Momentum Joint ShrinkingWenshuo Li, Xinghao Chen, Han Shu, Yehui Tang 等ICML 2024 · 被引用 11 次
- Efficient Fault Tolerance for Recommendation Model Training via Erasure CodingTianyu Zhang, Kaige Liu, Jack Kosaian, Juncheng Yang 等VLDB 2023 · 被引用 10 次
- FM-Delta: Lossless Compression for Storing Massive Fine-tuned Foundation ModelsWanyi Ning, Jingyu Wang, Qi Qi, Mengde Zhu 等NeurIPS 2024 · 被引用 10 次
相关 Paper
- Indirect Stochastic Gradient Quantization and Its Application in Distributed Deep LearningAfshin Abdi, Faramarz FekriAAAI 2020 · 被引用 5 次
- Towards Lossless Memory-efficient Training of Spiking Neural Networks via Gradient Checkpointing and Spike CompressionYifan Huang, Wei Fang, Zecheng Hao, Zhengyu Ma 等ICLR 2026
- SK-Gradient: Efficient Communication for Distributed Machine Learning with Data SketchJie Gui, Yuchen Song, Zezhou Wang, Chenhong He 等ICDE 2023 · 被引用 9 次
- Lossy and Lossless (L2) Post-training Model Size CompressionYumeng Shi, Shihao Bai, Xiuying Wei, Ruihao Gong 等ICCV 2023 · 被引用 5 次
- Communication Efficient SGD via Gradient Sampling With Bayes PriorLiuyihan Song, Kang Zhao, Pan Pan, Yu Liu 等CVPR 2021
