IncrCP: Decomposing and Orchestrating Incremental Checkpoints for Effective Recommendation Model Training
Qingyin Lin, Jiangsu Du, Rui Li, Zhiguang Chen, Wenguang Chen, Nong Xiao
Abstract
Training large models for modern recommendation systems requires a substantial number of computational devices and extended periods. Since it is essential to store model checkpoints throughout the training progress for accuracy debugging or mitigating potential failures, checkpointing systems are widely used. However, given that recommendation models can scale to hundreds of gigabytes or more, existing solutions often introduce significant overhead in terms of both storage and I/O. In this paper, we present IncrCP, a checkpointing system specifically designed for recommendation models. Given that only a small fraction of model parameters are modified in each iteration, IncrCP creatively leverages the incremental checkpointing strategy and overcomes the inherent slow recovery problem. To support recovering all states throughout the training process, while also ensuring efficient storage utilization and rapid recovery, IncrCP proposes the 2-D chunk approach. It proactively records changed parameters in the training process as well as their indexes, extracts parameters according to duplicated indexes as independent chunk files, and then orchestrates these chunks in the 2-dimensional linked list. In this way, IncrCP achieves fast recovery by loading less unnecessary parameters and performing less deduplication during recovery. Furthermore, IncrCP includes a selective extraction approach to reduce I/O by avoiding worthless extractions and a concatenate approach to reduce random disk access when recovery. Evaluations show that IncrCP achieves up to 6.6× recovery speedup compared to the naive incremental strategy and saves storage space by 60.4% with slight overhead compared to another recovery-friendly strategy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0dfd45ba-6696-45d6-b0ae-e27d8d6ad666Builds on8
- CheckFreq: Frequent, Fine-Grained DNN CheckpointingJayashree Mohan, Amar Phanishayee, Vijay ChidambaramFAST 2021 · 175 citations
- GEMINI: Fast Failure Recovery in Distributed Training with In-Memory CheckpointsZhuang Wang, Zhen Jia, Shuai Zheng, Zhen Zhang et al.SOSP 2023 · 61 citations
- Bagpipe: Accelerating Deep Recommendation Model TrainingSaurabh Agarwal, Chengpo Yan, Ziyi Zhang, Shivaram VenkataramanSOSP 2023 · 17 citations
- EL-Rec: Efficient Large-Scale Recommendation Model Training via Tensor-Train Embedding TableZheng Wang, Yuke Wang, Boyuan Feng, Dheevatsa Mudigere et al.SC 2022 · 15 citations
- GPU-Enabled Asynchronous Multi-level Checkpoint Caching and PrefetchingAvinash Maurya, M. Mustafa Rafique, Thierry Tonellot, Hussain J. AlSalem et al.HPDC 2023 · 13 citations
Related papers
- LowDiff: Efficient Frequent Checkpointing via Low-Cost Differential for High-Performance Distributed Training SystemsChenxuan Yao, Feifan Liu, Yuchong Hu, Zhengyu Liu et al.SC 2025 · 3 citations
- Check-N-Run: a Checkpointing System for Training Deep Learning Recommendation ModelsAssaf Eisenman, Kiran Kumar Matam, Steven Ingram, Dheevatsa Mudigere et al.NSDI 2022
- Efficient Fault Tolerance for Recommendation Model Training via Erasure CodingTianyu Zhang, Kaige Liu, Jack Kosaian, Juncheng Yang et al.VLDB 2023 · 10 citations
- PCcheck: Persistent Concurrent Checkpointing for MLFoteini Strati, Michal Friedman, Ana KlimovicASPLOS 2025 · 11 citations
- Kraken: memory-efficient continual learning for large-scale real-time recommendationsMinhui Xie, Kai Ren, Youyou Lu, Guangxu Yang et al.SC 2020 · 38 citations
