ExCP: Extreme LLM Checkpoint Compression via Weight-Momentum Joint Shrinking
Wenshuo Li, Xinghao Chen, Han Shu, Yehui Tang, Yunhe Wang
摘要
Large language models (LLM) have recently attracted significant attention in the field of artificial intelligence. However, the training process of these models poses significant challenges in terms of computational and storage capacities, thus compressing checkpoints has become an urgent problem. In this paper, we propose a novel Extreme Checkpoint Compression (ExCP) framework, which significantly reduces the required storage of training checkpoints while achieving nearly lossless performance. We first calculate the residuals of adjacent checkpoints to obtain the essential but sparse information for higher compression ratio. To further excavate the redundancy parameters in checkpoints, we then propose a weight-momentum joint shrinking method to utilize another important information during the model optimization, i.e., momentum. In particular, we exploit the information of both model and optimizer to discard as many parameters as possible while preserving critical information to ensure optimal performance. Furthermore, we utilize non-uniform quantization to further compress the storage of checkpoints. We extensively evaluate our proposed ExCP framework on several models ranging from 410M to 7B parameters and demonstrate significant storage reduction while maintaining strong performance. For instance, we achieve approximately compression for the Pythia-410M model, with the final performance being as accurate as the original model on various downstream tasks. Codes will be available at https://github.com/Gaffey/ExCP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- QT-ViT: Improving Linear Attention in ViT with Quadratic Taylor ExpansionYixing Xu, Chao Li, Dong Li, Xiao Sheng 等NeurIPS 2024 · 被引用 7 次
- IncrCP: Decomposing and Orchestrating Incremental Checkpoints for Effective Recommendation Model TrainingQingyin Lin, Jiangsu Du, Rui Li, Zhiguang Chen 等VLDB 2025 · 被引用 2 次
- CCL-D: A High-Precision Diagnostic System for Slow and Hang Anomalies in Large-Scale Model TrainingYida Gu, Fakang Wang, Jianhao Fu, Zhenhang Sun 等PPoPP 2026
它引用的顶会 Paper6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy 等ICLR 2020 · 被引用 1,037 次
- RMSprop converges with proper hyper-parameterNaichen Shi, Dawei Li, Mingyi Hong, Ruoyu SunICLR 2021 · 被引用 79 次
相关 Paper
- Compressing Large Language Models by Joint Sparsification and QuantizationJinyang Guo, Jianyu Wu, Zining Wang, Jiaheng Liu 等ICML 2024 · 被引用 33 次
- On Efficient Constructions of CheckpointsYu Chen, Zhenming Liu, Bin Ren, Xin JinICML 2020 · 被引用 21 次
- ACBQ: Adaptive Cross-Block Quantization of Large Language ModelsHailing Wang, Jianglin Lu, Yitian Zhang, Huimin Zeng 等ACL 2026
- ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM TrainingWenxiang Lin, Xinglin Pan, Ruibo Fan, Shaohuai Shi 等SIGCOMM 2026 · 被引用 1 次
- AdaCheck: An Adaptive Checkpointing System for Efficient LLM Training with Redundancy UtilizationWeijie Liu, Shengwei Li, Zhiquan Lai, Keshi Ge 等FAST 2026 · 被引用 3 次
